金融質問の曖昧さを確認して解くAIの能力を評価
FinInteract: Benchmarking Clarification and Intent Integration in Ambiguous Financial Question Answering
この論文をやさしく読む
ひとことで言うと
正しい資料を見つけても、質問者が会社全体と一部門のどちらを指すかを間違えると答えはずれます。この研究は、AIが自分から必要な確認をし、返答を回答に反映できるかを評価します。
何に役立つ?
金融資料を読むAIの評価に、数値の照合だけでなく質問意図の確認を加えるために役立ちます。提示された営業利益額は要旨の曖昧さの例であり、ここで最新の決算値として確認したものではありません。
この研究の面白いところ
同じ回答でも、評価側が一般的な解釈を正解にすると正答率が3.1倍になりました。解釈を教えれば90%超でも、自分で引き出す条件では最高28.9%という能力差を可視化しています。
どこまで分かった?
対象は英中二言語の173事例です。五分類すべてに同じ強さの統計的根拠があるわけではなく、対象範囲と指標定義以外は探索的な評価です。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
大規模言語モデルのエージェントは、規制当局への提出資料を検索して金融関連の質問に答える機会が増えている。この種の質問は、一見明確でも必要な条件が不足していることが多い。例えばMeta Platformsの「営業利益」は連結では467.5億ドルだが、Family of Apps部門では628.7億ドルであり、いずれの解釈も提出資料に照らして正確に確認できる。能力のあるエージェントは、もっともらしくても意図されていない解釈に決め打ちせず、曖昧さに気付いて質問するべきである。既存の金融ベンチマークはこれを測れない。質問ごとに正解が一つしかないため、曖昧さを解消するエージェントと、よくある解釈を推測したエージェントを区別できないからである。この盲点を「単一正解の錯覚」と呼ぶ。 本研究は、英語と中国語の二言語からなる173事例のベンチマークFinInteractを公開する。曖昧さを五つに分類し、各質問に既定の解釈と意図された解釈を対応付け、エージェントが適切な確認情報を引き出し、その後それを取り込むかを採点する。同一の出力を意図された解釈ではなく既定の解釈に対して再採点すると、GPT-4oの正答率は3.1倍に膨らみ、この錯覚が確認された。 さらに、解釈を与えればモデルは90%を超える正答率を示す一方、自ら確認して解釈を引き出す必要がある場合は最高でも28.9%だった。適切な確認対象を捉える能力は曖昧さの分類間で一様ではない。この分類では、企業・事業等の対象範囲と指標定義については十分な検出力があるが、その他の分類の評価は探索的である。また、曖昧さの分類を条件として与えると、推論時にも学習時にも曖昧さの解消が改善した。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-21(UTC)
- 最新改訂
- 2026-09-21 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-21 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Large language model agents increasingly answer financial questions by searching regulatory filings. Such questions are often deceptively under-specified: Meta Platforms' "operating income" is $46.75B consolidated but $62.87B for the Family of Apps segment, and each reading is exactly verifiable against the filing. A capable agent should recognize the ambiguity and ask, rather than commit to a plausible but unintended reading. Existing financial benchmarks cannot measure this, because one gold answer per question cannot separate agents that resolve the ambiguity from those that guess the common reading, a blind spot we call the single-gold illusion. We release FinInteract, a bilingual (English/Chinese) benchmark of 173 instances that pairs each question with a default and an intended interpretation across a five-category ambiguity taxonomy, and grades whether an agent elicits the right clarification and then integrates it. Re-grading identical outputs against the default rather than the intended reading inflates GPT-4o's accuracy by 3.1 times, confirming the illusion. Beyond it, we find that models answer above 90% once the interpretation is supplied but at most 28.9% when they must elicit it themselves, that targeting is uneven across a taxonomy well powered for entity scope and metric definition and exploratory elsewhere, and that conditioning on the ambiguity category improves resolution at both inference and training time.
arXiv ID: 2609.24002 / 要約の誤りについて