arXiv論文メモ
新着一覧
cs.CL / cs.CR / cs.IR · 査読状況未確認

ベトナムの個人データ処理記録をローカルAIで抽出

Automated Extraction of Records of Processing Activities (RoPA) Using Hybrid RAG and Locally Deployed Large Language Models

To Duy Hinh, Nguyen Le Quoc Anh, Phan Van Tri, and Khuong Nguyen-An

この論文をやさしく読む

ひとことで言うと

個人データ処理の記録に必要な情報を、組織内で動く言語モデルで文書から取り出す方法です。

何に役立つ?

RoPAの作成支援で、クラウドへ情報を送らない構成を検討する際に参考となる。法律への適合を自動的に保証する結果ではない。

この研究の面白いところ

採点器のF1、抽出の網羅率、専門家の審査結果を区別して報告し、採点器の高いF1を抽出精度と混同していない。

どこまで分かった?

抽出のトークン網羅率は50.04〜55.25%で、値単位の適合率は未測定。専門家審査は参照値の35.9%を対象とした。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

ベトナムの個人データ保護法(法律91/2025/QH15)と政令356/2025/ND-CPは、2026年1月1日から、組織に個人データ処理活動の記録(RoPA)の作成と維持を求める。手作業での準備は労力が大きく、クラウド上の大規模言語モデルはデータ主権の要件と衝突する可能性がある。著者らは、tsvectorによる語彙順位付け、密なベクトル検索、Reciprocal Rank Fusion(RRF)、ローカル配置した言語モデルを組み合わせた検索で、RoPA情報を自動抽出するRoPA Managerを提案する。32組織、77の処理活動、12の項目群、4,338の参照値からなるベトナム語のRoPAベンチマークを導入し、三つの異なる段階で評価した。言語モデルを使わず、値を攪乱したデータで試した自動採点器はF1=0.9493[0.9436、0.9548]を達成した。これは採点器の頑健性を測る値であり、端から端までの抽出精度ではない。全体の抽出では、参照ラベルに対するトークンの網羅率は50.04〜55.25%だった。独立した二人の専門家は、参照値1,558件、ベンチマークの35.9%を審査し、誤った値は見つけず、一致率99.68%、PABAK=0.9936を得た。ただし値単位の適合率は測っていない。24 GBのGPU上で32組の対応する条件を比較すると、ローカル配置のQwen3.5-27B-GPTQ-Int4とクラウドのDeepSeek-V4-Flashに統計的に有意な差はなかった。差はDeepSeekに有利な0.20ポイントで、95%信頼区間は−0.93〜1.32、p=0.72だった。一方、Gemma-4-31Bは有意に劣った(p<0.01)。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-23(UTC)
最新改訂
2026-09-23 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Vietnam's Personal Data Protection Law (Law No. 91/2025/QH15) and Decree No. 356/2025/ND-CP, effective January 1, 2026, require organizations to establish and maintain Records of Processing Activities (RoPA). Manual RoPA preparation is labor-intensive, while cloud-hosted large language models (LLMs) may conflict with data-sovereignty requirements. We propose RoPA Manager, a system for automated RoPA information extraction using hybrid retrieval that combines lexical ranking over tsvector, dense-vector search, Reciprocal Rank Fusion (RRF), and locally deployed LLMs. We introduce a Vietnamese RoPA benchmark with 32 organizations, 77 processing activities, 12 field groups, and 4,338 reference values. Evaluation is reported at three distinct levels. The automated scorer, tested on perturbed data without invoking an LLM, achieved F1 = 0.9493 [0.9436, 0.9548]; this measures scorer robustness rather than end-to-end extraction accuracy. End-to-end extraction achieved token coverage of 50.04-55.25% against the reference labels. Two independent experts reviewed 1,558 reference values (35.9% of the benchmark), found no incorrect values, and achieved 99.68% agreement with PABAK = 0.9936. Value-level precision was not measured. Across 32 paired scenarios on a 24 GB GPU, locally deployed Qwen3.5-27B-GPTQ-Int4 showed no statistically significant difference from cloud-based DeepSeek-V4-Flash (difference 0.20 percentage points in favor of DeepSeek, 95% CI [-0.93, 1.32], p = 0.72), while Gemma-4-31B performed significantly worse (p < 0.01).

著者のコメント

English version followed by Vietnamese version. Accepted for publication in the Proceedings of the 29th National Conference on Selected Issues of Information and Communication Technology (VNICT 2026), Hanoi, Vietnam, November 7-8, 2026

arXiv ID: 2609.27359 / 要約の誤りについて