arXiv論文メモ
新着一覧
cs.CV · 査読状況未確認

製品画像と地域別法令を結ぶたばこ規制の検索手法

PRISM-RAG: Multimodal Hypergraph Retrieval-Augmented Generation for Tobacco Product and Legislative Policy Reasoning

Manuel Serna-Aguilera, Raegan Anderes, Page Dobbs, Khoa Luu

この論文をやさしく読む

ひとことで言うと

製品の画像を手掛かりに法令を探すとき、似た内容でも別の地域の法律を混ぜないようにする検索方法です。

何に役立つ?

考えられる用途は、製品と地域を指定して規制文書を調べる作業の支援です。要旨の結果は構築したデータセットでの評価です。

この研究の面白いところ

内容の類似度だけでなく法域を検索経路に組み込み、索引作成にLLMを使わずに精度を改善しています。

どこまで分かった?

93.9%は正しい法域の文章を取得した割合で、法的判断全体の正解率ではありません。改善幅は48.6%ではなく48.6パーセントポイントです。対象は米国13法域です。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

法域をまたいで意味の似た法令文を区別することは、既存手法では解決できていない検索の問題である。この文脈間の衝突により、生成モデルは話題としては関連していても法域が誤った情報源に基づき、確信を持った回答を生成しかねない。たばことニコチンの規制は米国内の法域によって異なる一方、よく似た表現を共有することが多い。そのため、頑健な推論には、単に関連文を検索するのではなく、ある製品にどの法域の法律が適用されるかを特定する必要がある。ポーチなどの新興製品は、曖昧な定義を利用して規制を回避する。最先端の文書検索拡張生成(RAG)手法は、この文脈間の衝突への対応が難しく、豊富な属性を記述したキャプションなどの画像属性を、類似する法令文群に結びつけることにも苦戦する。 本研究ではNicoPRISM(Nicotine Product and Regulation Image-and-Text Surveillance Multimodal)を導入する。これは161,563枚の画像、属性キャプション、米国13法域を対象とする製品・健康・法令文書の知識基盤、そして政策適合性QAと製品知識QAの2課題にわたる、検証済みの質問回答1,495組からなる。 さらに、索引作成時にLLMを一切呼び出さず、画像・キャプション・実体の上に構築するマルチモーダルなハイパーグラフRAGの枠組みPRISM-RAGを提案する。各問い合わせを製品画像に基づかせ、法域を考慮した文脈構成機構を通じて検索を振り分けることで、問い合わせ対象の法域の法令文が言語モデルに届くことを構成上保証する。PRISM-RAGは政策適合性の問い合わせの93.9%で正しい法域の文章を取得し、標準的なRAGに対して48.6パーセントポイントの優位性を示した(p<0.001)。索引作成時のLLM呼び出しはゼロ、問い合わせ時は1回であり、キーワード、意味、法域の正確さ、適合性の正確さという指標全体で、最先端のRAGと同等以上の性能を示した。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-20(UTC)
最新改訂
2026-09-20 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

The disambiguation of semantically similar statutory text across jurisdictions is a retrieval problem that existing methods do not solve. This inter-context conflict can steer generative models toward confidently produced answers grounded in topically relevant but jurisdictionally incorrect sources. Tobacco and nicotine regulations vary by US jurisdiction, often sharing similar language, thus, robust reasoning requires identifying which jurisdiction's law governs a given product, not merely retrieving relevant text. Emerging products (e.g., pouches) exploit ambiguous definitions to evade regulation. State-of-the-art (SOTA) document retrieval-augmented generation (RAG) methods struggle to address this inter-context conflict, and thus struggle to connect image attributes (e.g., rich attribute captions) to the set of similar legislation texts. We introduce NicoPRISM (Nicotine Product and Regulation Image-and-Text Surveillance Multimodal), comprising 161,563 images, attribute captions, a knowledge base of product, health, and legislative documents spanning 13 US jurisdictions, and 1,495 validated question-answer pairs across two tasks: policy compliance QA and product knowledge QA. We also propose PRISM-RAG, a multimodal hypergraph RAG framework built over images, captions, and entities without any LLM calls at index time, grounding every query in a product image and routes retrieval through a jurisdiction-aware context assembly mechanism guaranteeing that statutory text from the queried jurisdiction reaches the language model by construction. PRISM-RAG retrieves passages from the correct jurisdiction in 93.9% of policy compliance queries, a 48.6 percentage point advantage over standard RAG (p<0.001), using zero LLM calls at index time and one at query time, and is competitive with or outperforms SOTA RAG frameworks across keyword, semantic, jurisdiction-, and compliance-accuracy metrics.

arXiv ID: 2609.23769 / 要約の誤りについて