遠くの文脈を取り出せても熟語の解釈が変わらない場合
MWE-ECL: Recoverable Long-Range Context Does Not Always Override Local Lexical Priors
この論文をやさしく読む
ひとことで言うと
長い文脈の情報を検索できても、熟語の意味判断に反映できないことがあるかを調べた。
何に役立つ?
長文を扱う言語モデルで、情報検索と回答への反映を分けて評価するのに役立つ。
この研究の面白いところ
手掛かりの検索、手掛かりなしの既定解釈、最終判断を対応づけて測る。
どこまで分かった?
モデルや言語によって差は異なり、中国語の一部では検索自体が不完全なため解釈への統合だけが原因とは言えない。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
長い文脈の評価では、モデルが遠くにある証拠を取り出せるかを試すことが多い。しかし、取り出せることが行動に影響するとは限らない。本研究は、離れた談話上の手掛かりを明示的に取り出せても、なじみのある複数語表現について、近くの語句から好まれる解釈を変えられない場合があるという予測を検証する。そのような失敗は、手掛かりがないときのモデルの既定の解釈と手掛かりが衝突する場合に集中し、もともと正しい判断はおおむね保たれるはずである。英語と中国語を扱う診断法Multiword Expression Effective Context Length(MWE-ECL)を導入する。対応をそろえた手掛かりの検索、手掛かりなしの既定解釈、解釈の三種類の質問によって、明示的な取り出しやすさ、モデルの既定傾向、手掛かりを条件とする判断をそれぞれ測る。 共通の0~128Kの文脈長で評価した英語の八つの運用条件では、既定解釈と手掛かりが衝突する項目の検索対照の正確さは0.989~1.000、既定解釈の上書きは0.806~1.000、正しく検索した場合に限ると0.809~1.000だった。もともと正しい判断の維持率は0.977~1.000だった。同じ呼び出しで検索と解釈を問う対照でも、DeepSeek V4 Proでは検索1.000に対し解釈0.900~0.920という差が再現された。したがって別々の呼び出しだけが理由ではない。他の二モデルでは差が小さいか存在せず、一般化できる範囲には限りがある。DeepSeek V4 Flashの別の質問適合試験では、512Kと1Mでも検索は完全だったが解釈は低く、紛らわしい候補と整合する手掛かりは検索よりも手掛かりなしの既定傾向を大きく変えた。この効果はモデル間で一様ではない。別途報告した中国語の10系列の部分集合にも似た記述的な差が見られたが、一部モデルの検索が不完全なため、統合だけの問題とは断定できない。MWE-ECLは、明示的に取り出せる遠方の文脈が、競合する近傍の意味判断を変えるかを評価する。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-23(UTC)
- 最新改訂
- 2026-09-23 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-23 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Long-context evaluations often test whether a model can recover distant evidence, but recoverability does not guarantee behavioral influence. We test the prediction that a distant discourse anchor can remain explicitly recoverable yet fail to change the locally preferred reading of a familiar multiword expression; such failures should concentrate when the model's no-anchor default conflicts with the anchor, while prior-correct decisions remain largely preserved. We introduce Multiword Expression Effective Context Length (MWE-ECL), a bilingual diagnostic whose matched anchor-retrieval, no-anchor prior, and interpretation prompts measure explicit recoverability, model-observed defaults, and anchor-conditioned decisions, respectively. Across eight English deployment panels on a shared 0-128K grid, retrieval-control accuracy on prior-conflict items is 0.989-1.000, prior-conflict override spans 0.806-1.000 (0.809-1.000 after conditioning on correct retrieval), and preservation of prior-correct decisions remains 0.977-1.000. A same-call control querying retrieval and interpretation in one prompt reproduces the gap for DeepSeek V4 Pro (1.000 retrieval versus 0.900-0.920 interpretation), showing that separate invocations are not its sole explanation; smaller or absent gaps in the other two models bound its generality. For DeepSeek V4 Flash, separate prompt-fit tests retain perfect retrieval with lower interpretation at 512K and 1M, while foil-consistent cues shift the no-anchor prior far more than retrieval; cross-model cue effects are heterogeneous. A separately reported 10-family Chinese subset shows similar descriptive gaps, but imperfect retrieval for some models prevents an integration-only attribution. MWE-ECL therefore evaluates whether explicitly recoverable distant context changes a competing local semantic decision.
arXiv ID: 2609.27590 / 要約の誤りについて