事実検証の得点向上を回答と証拠に分けて調べる
What Changes When Fact-Verification Scores Improve? Evidence and Answer Accounting Across Trained Verifiers and LLMs
この論文をやさしく読む
ひとことで言うと
事実検証の得点が上がった理由を、回答の改善と提出証拠の改善に分けて測る研究。
何に役立つ?
事実検証システムの比較で、得点向上が回答の正しさと証拠の質のどちらから来たかを判断するのに役立つ。
この研究の面白いところ
回答を固定し、採点器に渡す証拠だけを替えて比較した。厳格な得点の9.61ポイントの上昇に対し、回答正解率の上昇は1.96ポイントだった。
どこまで分かった?
95%区間は指定の4つのチェックポイントを条件とする。文脈量による効果は事前指定のデータセット横断基準には達していない。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
事実検証の総合得点は、回答と提出された証拠を一緒に評価する。得点が上がったとき、回答を固定しても改善はどれだけ残るのか。FEVEROUSの厳格な得点は、正しい回答があり、提出した証拠に注釈付きの完全な証拠群が含まれる主張の割合である。学習済みDeBERTaの4つのチェックポイントと7,890件の主張について、DCUFの証拠をUnifEEの証拠に替えると、厳格な得点は9.61ポイント上昇したが、回答の正解率の上昇は1.96ポイントだった。厳格な得点の上昇について、対応のある比較による95%区間は、この4つのチェックポイントを条件として8.77~10.43ポイントである。採点器に渡す証拠だけを替え、DCUFまたはUnifEEの証拠から生成した回答をそれぞれ固定すると、改善のうち7.92または9.08ポイントを説明できた。 この証拠による改善が評価方法にどう依存するかを調べるため、2つの80億パラメータの言語モデルを用い、FEVER、FEVEROUS、SciFactで、2種類の回答形式と2種類の文脈量を組み合わせて470,400件の応答を生成した。文脈を256トークンから2,048トークンに増やすと、FEVEROUSで回答を固定した際の証拠による得点改善は、Qwenで3.84ポイント、Llamaで3.10ポイント増えた。その効果は事前に定めたデータセット横断の判定基準に届かなかった一方、一部の区間は2ポイントという小効果の境界を越えた。事後分析では回答と提出証拠の変化を定量化し、全体の正解率や証拠の網羅率だけでは主張ごとの傾向を捉えられない場合を示した。回答と証拠の4通りの組み合わせによる得点は、最終得点と集計指標だけでは見えない変化を明らかにする。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-22(UTC)
- 最新改訂
- 2026-09-22 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-22 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
A joint fact-verification score assesses answers and submitted evidence together. When the score improves, how much of the gain remains if the answers are held fixed? On FEVEROUS, strict score is the percentage of claims with a correct answer and a complete annotated evidence group in the submitted evidence. Across four trained DeBERTa checkpoints and 7,890 claims, replacing DCUF evidence with UnifEE evidence raises strict score by 9.61 percentage points, compared with 1.96 percentage points in answer accuracy. The paired 95% interval for the strict-score gain is [8.77, 10.43], conditional on these checkpoints. Replacing only the evidence passed to the scorer accounts for 7.92 or 9.08 percentage points when we retain the answers generated from DCUF or UnifEE evidence, respectively. To examine how this evidence gain depends on evaluation choices, we generate 470,400 responses from two 8B LLMs on FEVER, FEVEROUS, and SciFact under two answer formats and two context budgets. Increasing context from 256 to 2,048 tokens raises the fixed-answer evidence gain on FEVEROUS by 3.84 and 3.10 percentage points for Qwen and Llama, respectively. The effects fall short of the prespecified cross-dataset criterion, while some intervals extend beyond the two-point small-effect bound. Post-hoc analyses quantify changes in answers and submitted evidence, and show when aggregate accuracy and evidence-coverage rates miss the claim-level pattern. The four answer-evidence score combinations reveal changes that endpoint and aggregate metrics leave unresolved.
著者のコメント
24 pages. Both authors contributed equally
arXiv ID: 2609.27064 / 要約の誤りについて