arXiv論文メモ
新着一覧
cs.AI · 査読状況未確認

規範文書の版と適用範囲を確認する質問応答を実運用評価

Version- and Scope-Aware Question Answering over Normative Documents: A Deployed System and an End-to-End Evaluation at Production Scale

Liuyin Wang, Shuaipeng Jin, Jiwei Shi, Jensen Hsu

この論文をやさしく読む

ひとことで言うと

規則を答えるAIが、検索で見つけた文章だけでなく、その版が有効か、質問者や地域に適用されるかを確認する仕組みです。生成の前に明示的なルールを使う方式を、ホスト型検索サービスと比較しています。

何に役立つ?

規程や制度を扱う質問応答で、古い版や対象外の規則を根拠にしないための設計参考になります。公開された質問、回答、採点、スクリプトは、評価の再現性を確認する材料になります。

この研究の面白いところ

文書量は約73,000件ですが、評価は正解の出典を持つ層化抽出200問で行っています。システム結果を見ずに質問を選ぶ規則を公開し、版・適用範囲の管理を評価可能にしています。

どこまで分かった?

97.7と88.1は総合スコアであり、要旨は正答率や法的な正確性の保証とは述べていません。登録者数と呼び出し数は運用規模の情報で、すべての回答が正しいことの証明ではありません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

規範文書に基づく質問へ正しく答えるには、一つの文章箇所の外にある情報が必要なことが多い。検索した文書が現在有効な版か、問題となる管轄・対象(機関や申請者など)・日付に適用されるか、そして規範に関する各主張を根拠となる原文へたどれるか、という情報である。ホスト型の検索サービスは、このような文書群を使う最初のシステムを作る開発コストを大幅に下げ、「文書をアップロードして質問する」方法を一般的な初期選択にした。 本研究では、実運用環境へ提供された約73,000件の規範文書候補を対象に、この初期選択を評価する。公開したベンチマークから層化抽出した200問を使い、すべての質問に正解の出典文書を用意する。公開した抽出規則は、システムの出力やスコアを一切参照しない。ホスト型サービスを、生成前に明示的な規則で版と適用範囲を解決する管理されたシステムと比較する。 管理されたシステムの総合スコアは97.7、ホスト型サービスは88.1であり、丸める前の平均から算出した差は9.6ポイントである。質問群、両システムで評価した回答文、スコア、報告されたベンチマーク統計を再現するスクリプトを公開している。 管理された構成は2026年1月から商用製品として稼働し、登録利用者は1,126人である。顧客組織にはZhipu AIとLecheng Healthが含まれる。2026年4月中旬までに、稼働日1日当たりの呼び出しは約10万回に達した。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-16(UTC)
最新改訂
2026-09-16 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Correctly answering a question grounded in normative documents often depends on information outside any single passage: whether the retrieved document is the version currently in force; whether it applies to the jurisdiction, subject (such as an institution or applicant), and date at issue; and whether each normative claim can be traced to its supporting source text. Hosted retrieval services have substantially lowered the engineering cost of building an initial system over such corpora, making "upload the documents and ask" a common default. We evaluate this default on approximately 73,000 candidate normative documents supplied to a production deployment. The evaluation uses a stratified sample of 200 questions from our published benchmark, with a gold source document for every question; the released sampling rule reads no system outputs or scores. We compare the hosted service with a governed system that resolves version and scope through explicit rules before generation. The governed system scored 97.7 overall, while the hosted service scored 88.1, a gap of 9.6 points computed from unrounded means. The question set, the answer text evaluated for both systems, the scores, and the scripts used to reproduce the reported benchmark statistics are public. The governed configuration has operated as a commercial product since January 2026 and serves 1,126 registered users; named customer organizations include Zhipu AI and Lecheng Health. By mid-April 2026, it had reached roughly 100,000 calls per workday.

著者のコメント

9 pages, 1 figure, 3 tables

arXiv ID: 2609.18769 / 要約の誤りについて