arXiv論文メモ
新着一覧
cs.CL · 査読状況未確認

金融情報を探すAIの根拠まで評価するFinFIRST

FinFIRST: Benchmarking Search Agents for Financial Information Retrieval, Sourcing and Traceability

Wenqing Wang, Haitao Xiang, Xinyi Zhao, Mingming Yin, Ying Zhong, Zhaoxin Huan, Qiheng Zhou, Jin Zhu, Xiaolu Zhang, Shi Chang, Jun Zhou

この論文をやさしく読む

ひとことで言うと

金融の質問に答えるAIについて、答えだけでなく情報源と計算過程も採点するベンチマークです。

何に役立つ?

金融情報検索エージェントの誤りが、取得・情報源の確認・計算のどこにあるかを調べる用途に役立ちます。個別の投資判断を検証するものではありません。

この研究の面白いところ

15構成を同じツール環境で比較し、計算と回答作成の成績が生情報の取得より低いという共通点を見つけています。

どこまで分かった?

結果は123課題、登録された138情報源、統一したツール環境での評価です。ほかの条件での順位までは要旨から分かりません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

金融情報の検索はLLMエージェントにとって要求の高い作業であり、最終回答の正しさだけでなく、時点に合った情報取得、権威ある情報源の選択、企業などの対象と期間の整合、単位と定義の一貫性、すべての結論に対する検証可能な根拠が必要となる。既存のベンチマークの多くは最終回答のみを評価するため、誤りの箇所を特定したり、回答に十分な根拠があるかを評価したりしにくい。この不足を補うため、著者らは細分化された評価基準を通じて回答とそれを支える証拠を合わせて評価する、初の金融ベンチマークFinFIRSTを提示する。FinFIRSTは専門家が作成した難易度の異なる123件の課題からなり、現実の金融場面に共通するパターンを基に、18項目の分類、6軸の範囲設計、138の金融情報源の登録簿、50人を超える金融専門家の貢献、6段階の品質管理を用いて構築した。各課題には根拠付きの参照資料があり、生情報の取得、情報源の検証、計算と回答の作成という3側面にわたる個別の基準に分解されている。統一したツール環境で15種類のモデル構成を評価した結果、Claude-Opus-5が個別基準の最高得点87.59%を、GPT-5.6-Solが厳格な合格率の最高値71.54%を得た。どのシステムでも、計算と回答の作成は生情報の取得より一貫して成績が低かった。FinFIRSTは最終回答の正確さを主目的に保ちながら、それを支える調査過程を測定、検証、診断できるようにする。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-21(UTC)
最新改訂
2026-09-21 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Financial search is a highly demanding task for LLM agents, requiring not only a correct final answer but also temporally valid information retrieval, authoritative source selection, entity and period alignment, unit and definition consistency, and verifiable evidence for all conclusions. Existing benchmarks predominantly evaluate only the final answer, making it difficult to localize errors or assess whether an answer is well-founded. To address this gap, we introduce FinFIRST (Financial Information Retrieval, Sourcing and Traceability), the first financial benchmark to jointly evaluate answers and supporting evidence through atomic rubrics. FinFIRST comprises 123 expert-authored tasks spanning a graduated difficulty spectrum, constructed from aggregate patterns of real-world financial scenarios through an 18-field taxonomy, a six-axis coverage blueprint, a registry of 138 financial sources, contributions from over 50 finance experts, and a six-stage quality-control pipeline. Each task is accompanied by an evidence-grounded reference package decomposed into atomic criteria across three dimensions: raw-information acquisition, source verification, and computation and answer formation. We evaluate 15 model configurations under a unified tool setting. Claude-Opus-5 achieves the highest atomic score of 87.59%, while GPT-5.6-Sol attains the highest strict pass rate of 71.54%. Computation and answer formation consistently lag behind raw-information acquisition across systems. FinFIRST retains final-answer correctness as the primary objective while making the supporting research process measurable, verifiable, and diagnosable.

著者のコメント

20 pages, 3 figures, and 7 tables. Dataset available at https://huggingface.co/datasets/inclusionAI/FinFIRST

arXiv ID: 2609.25192 / 要約の誤りについて