LLMの要因説明を正直にさせる監査の選び方
How the Audit Rule Shapes Faithful Factor Explanations in LLMs
この論文をやさしく読む
ひとことで言うと
AIが「どの情報を判断に使ったか」を説明するとき、重要だと答えた項目だけ調べる監査は、影響を小さく報告する動機を生むという研究です。
何に役立つ?
限られた検証回数でAIの説明を点検する仕組みの設計に役立ちます。報告内容と無関係に調べる機会を一定程度残すことが提案されています。
この研究の面白いところ
採点方法が適正でも、何を監査するかの選び方によって正直な報告が不利になる点を示しています。理論、合成エージェント、実際のLLMの評価を組み合わせています。
どこまで分かった?
実際のLLMで理論と同じ誘因が観察されたのは、その誘因を明示した条件です。また、示されるのは全面的な報告抑制より正直な報告を好ましくする結果であり、あらゆる説明の忠実性を保証するものではありません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
大規模言語モデルは、自らの出力にどの入力要因が影響したかを尋ねられることが多い。構造化された入力であれば、このような報告は反実仮想的な摂動によって確認できる。しかし、効果を推定するには各要因について複数回問い合わせる必要があるため、通常は検証予算が限られる。本研究では、この予算制約が、要因ごとの影響を正直に報告する誘因をどう変えるかを調べる。 相互作用を検証ゲームとして定式化し、監査が報告に依存する場合には、適正なスコアリングだけでは不十分であることを示す。報告依存の監査は、重要だと報告した要因ほど確認されやすく、推定ノイズに対して罰を受けやすいため、影響の報告を抑える誘因を生む。これに対し、報告に依存しない監査、または小さくても報告非依存の監査確率の下限を持つ混合規則は、この経路を取り除き、すべての影響を報告しないことよりも正直な報告を望ましいものにする。 Counterfactual Brier Score(CBS)を使って枠組みを具体化し、4つの自然言語処理ベンチマークで予測を評価する。合成された合理的エージェントは理論的予測と厳密に一致し、実際のLLMも誘因を明示すると同じ誘因に従う。設計上の主要な示唆は明快である。部分的な検証のもとでは、要因単位の説明システムに報告非依存の監査成分を含め、過少報告によって監視を免れられないようにすべきである。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-10-01(UTC)
- 最新改訂
- 2026-10-01 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-10-01 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Large language models are often asked which input factors influenced their outputs. For structured inputs, such reports can be checked by counterfactual perturbation, but each factor must be queried multiple times to estimate its effect, so verification is usually budget-limited. We study how this limited-budget setting changes the incentive to report factor-level influence truthfully. We formalize the interaction as a verification game and show that proper scoring alone is not enough when auditing depends on the report: report-dependent auditing creates a suppression incentive, because factors reported as important are more likely to be checked and penalized for estimation noise. In contrast, report-independent auditing, or a mixed rule with a small report-independent floor, removes this channel and makes truthful reporting preferable to full suppression. We instantiate the framework with the Counterfactual Brier Score (CBS) and evaluate its predictions on four NLP benchmarks. A synthetic rational agent matches the theoretical prediction exactly, and real LLMs follow the same incentives when they are made explicit. The main design implication is simple: under partial verification, factor-level explanation systems should include a report-independent audit component so that under-reporting cannot be used to avoid scrutiny.
arXiv ID: 2610.01514 / 要約の誤りについて