長文質問応答では学習型の証拠選択が強い検索を上回らない
When Learned Context Planning Fails to Beat Strong Retrieval: A Controlled Study of Planning, Routing, and Reranking for Long-Context QA
この論文をやさしく読む
ひとことで言うと
長文質問応答で、学習した証拠の選択方法を強い検索方法と比べた。
何に役立つ?
長文を扱うAIの設計で、複雑な計画器を追加する価値を評価するのに役立つ。
この研究の面白いところ
未使用の152問でも、アンカー付き検索が計画器利用法を42.11%対36.84%で上回った。
どこまで分かった?
503問の主分析は訓練・開発問題を含む。結論は使用モデル、データ、文字数予算などこの実験設定についてのもの。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
学習型の文脈計画は、回答モデルが推論する前に証拠となる要素を選ぶ。本研究は、長い文脈を使う多肢選択式質問応答で、強い検索、経路選択、予算付き選択器、再順位付けを対照として用いた後でも、この学習型選択が改善をもたらすかを試す。主な診断にはLongBench-v2の503問すべてとQwen2.5-7B-Instructを使用した。計画器は訓練140問と開発28問から、結果に基づいて選んだ処理履歴で教師あり微調整した。503問の分析にはそれらの問題も含まれるため、部分的に推移的な評価である。 1万8千文字の予算では、アンカー付きハイブリッド検索の正解率は36.18%、BM25は35.98%で、最良の直接的な計画器利用法は34.19%だった。手を付けていない152問のテスト分割でも、アンカー付きハイブリッド検索が42.11%対36.84%で上回った。漏洩を防ぐ経路選択器でも、大きな理想上限との差を埋められなかった。予算が厳しい場合、最良の計画器は6千文字で0.40ポイント上回るだけで、9千文字では下回った。計画器を使う再順位付けは6千文字で推定1.79ポイントの改善だったが、対応のある区間推定はゼロをまたぎ、9千文字では対照と同率だった。詰め込み順序とスコアの平坦さの分析からも安定した機構は見いだせなかった。この設定では学習型計画は弱い関連性の手掛かりであり、強い検索の代わりにはならない。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-22(UTC)
- 最新改訂
- 2026-09-22 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-22 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Learned context planning selects evidence atoms before an answer model reasons over them. We test whether this learned selection improves long-context multiple-choice QA after strong retrieval, routing, budgeted-selector, and reranking controls. Our primary diagnostic uses all 503 LongBench-v2 MCQ questions with Qwen2.5-7B-Instruct. The planner is SFT-trained on outcome-selected traces from 140 training and 28 development questions; because the 503-question analysis includes those questions, it is partly transductive. At an 18k-character budget, anchored hybrid retrieval reaches 36.18% accuracy and BM25 reaches 35.98%, while the best direct planner-guided method reaches 34.19%. On the untouched 152-question test split, anchored hybrid remains higher (42.11% versus 36.84%). Leakage-safe routers cannot convert a large oracle gap. Under tight budgets, the best planner is ahead by only 0.40 points at 6k and loses at 9k; planner-guided reranking has a +1.79-point estimate at 6k with a paired interval crossing zero and ties the control at 9k. Packing-order and score-flatness analyses did not identify a stable mechanism. Under this setup, learned planning is a weak relevance signal rather than a replacement for strong retrieval.
著者のコメント
5 pages. Accepted at the Seventh Workshop on Insights from Negative Results in NLP (Insights 2026), co-located with EMNLP 2026
arXiv ID: 2609.26976 / 要約の誤りについて