arXiv論文メモ
新着一覧
cs.CL / cs.AI · 査読状況未確認

紛らわしい文脈に左右されない文脈内学習の訓練法

When Context Misleads: In-context Learning with Jurisdiction in Large Language Models

Pei-lin Li, Qingle Liu, Junyang Feng, Siyu Li, Sunqi Fan, Xin-Sheng Chen, Shuojin Yang

この論文をやさしく読む

ひとことで言うと

例として示された情報を鵜呑みにせず、回答に使うべきか判断するようLLMを訓練する方法である。

何に役立つ?

誤った文脈が混じる場面で、モデルの回答を事実に沿わせる訓練や評価に役立つ。

この研究の面白いところ

文脈内学習の得点と、疑似科学的な文脈への耐性を同時に改善することを狙った。

どこまで分かった?

七分野のベンチマークと四つの基盤モデルでの結果である。あらゆる種類の誤情報に対する耐性を保証しない。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

文脈内学習(ICL)は、現在の大規模言語モデルの利用で中心的な役割を持つ。しかし既存のICL追加学習法には重大な見落としがある。例示からパターンを抽出することには優れる一方、その文脈情報を最終回答の根拠として採用すべきかを判断する「文脈の権限」を十分に扱わない。この能力を評価するため、七つの分野の疑似科学的主張を含むFakeContextBenchを導入する。商用モデルとオープンソースモデルの評価から、大規模な事前学習だけでは、文脈の権限を確実に判別できないことが分かった。さらに、一般的なICL微調整法は、紛らわしい文脈に影響されやすくし、基盤モデルと比べて現実に即した回答の正確さを最大14.95ポイント下げる場合があった。 この両立の課題に対し、文脈の検証を学習目標に取り入れる追加学習枠組みJurisdiction In-Context Learning(J-ICL)を提案する。四つの基盤モデルで、J-ICLは対応する基盤モデルと比べ、ICLEvalを平均5.84ポイント、現実に即した正確さを9.20ポイント改善した。また、MetaICLおよびSymbol Tuningと比べ、Reality Rateを平均18.09ポイント高めた。結果は、ICLの能力と欺く文脈への耐性を同時に改善できることを示す。ベンチマークは指定されたGitHubリポジトリで公開されている。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-23(UTC)
最新改訂
2026-09-23 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

In-Context Learning (ICL) has become a cornerstone of modern LLM deployment. However, existing ICL post-training methods have a critical blind spot: they excel at extracting patterns from demonstrations while often neglecting context authority, the ability to determine whether contextual information should govern the final answer. To benchmark this capability, we introduce FakeContextBench, which contains pseudoscientific claims across seven domains. Our evaluation of commercial and open-source models shows that large-scale pre-training alone is insufficient for reliable context-authority discrimination. Moreover, prevalent ICL fine-tuning methods can increase susceptibility to misleading context, reducing reality accuracy by up to 14.95 percentage points relative to the base model. To address this trade-off, we propose Jurisdiction In-Context Learning (J-ICL), a post-training framework that incorporates context validation into the training objective. Across four model backbones, J-ICL improves ICLEval by an average of 5.84 percentage points and reality accuracy by 9.20 points over the corresponding base models. It also raises the Reality Rate by an average of 18.09 points relative to MetaICL and Symbol Tuning. These results demonstrate that ICL capability and resistance to deceptive context can be improved together. The benchmark is available at https://github.com/peilin717/FakeContext-Bench.

arXiv ID: 2609.27603 / 要約の誤りについて