arXiv論文メモ
新着一覧
cs.LG / cs.AI · 査読状況未確認

自律エージェントの説明が誤る条件を調べる

Triggers and Diagnostics for LLM-Based Interpretability Failures in Active Inference Agents

Param Raval, Rohit Shenoy, Archana Vaidheeswaran

この論文をやさしく読む

ひとことで言うと

自律エージェントの行動を説明するLLMが、誤った状態や行動をもっともらしく説明するか調べる。

何に役立つ?

エージェントの運用監視に生成された説明を使う場合、その説明自体を検査する必要性を判断する材料になる。

この研究の面白いところ

観測値の改変、誤行動、メタデータへの文章注入で、説明の誤りをそれぞれ確認する。

どこまで分かった?

緩和策は提案のみで未評価。結果は指定の電力需給エージェント、三つの説明基盤、試験条件に基づく。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

自律エージェントの実行時監視に大規模言語モデル(LLM)による説明機能を付け、運用者が内部状態の代わりに、生成された信念や行動の説明を読むことが増えている。本研究は説明自体を監査する。ドイツの電力需要を追跡し発電量を調整する能動推論(AIF)エージェントと、GPT-4o、Claude-3-Opus、Geminiという三種類の基盤を使うLLM説明機能を組み合わせ、ブラックボックスの三種類の誘因で調べる。 観測データを一段階につき600 MWずつ改変すると、エージェントの事後推定は490 MW、送電網容量の約0.9%動く。この注入中に作られた30件の説明は、定めた評価基準ではどれも異常を指摘せず、改変された信念を流暢に語った。エージェントが客観的に誤った行動をした時点では、三種類の説明機能は80~95%の割合で追従的な正当化を生成した(各基盤20件)。観測メタデータの攻撃者が制御できる文章は説明機能を誘導し、影響されやすさは提供元によって異なるが、データの外部流出は三つすべてで成功した。 著者らは各失敗に対する緩和策を提案するが、その評価は行っていない。観察されたすべての失敗で、説明は流暢だが誤っていた。さらに、この説明機能の構造には、運用者が説明に基づいて行動する前に、その説明の真偽を検査する仕組みがない。したがって、エージェントを配備する際の監査には説明機能の試験も含めるべきだと論じる。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-19(UTC)
最新改訂
2026-09-19 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

LLM explainers are increasingly attached to autonomous agents as runtime oversight, with operators reading a generated account of the agent's beliefs and actions rather than its internal state. We audit the account itself, pairing an Active Inference (AIF) agent that tracks German grid demand and adjusts generation with an LLM explainer on three backends (GPT-4o, Claude-3-Opus, Gemini), and probing the pair with three black-box triggers. Corrupting the observation stream by 600 MW per step moves the agent's posterior by 490 MW, roughly 0.9% of grid capacity. None of the 30 explanations produced during the injection flag anything under a stated rubric, and each narrates the corrupted belief fluently. On timesteps where the agent takes an objectively wrong action, all three explainers produce a sycophantic rationalization 80-95% of the time (n = 20 per backend). Attacker-controlled text in the observation metadata field steers the explainer, with susceptibility differing by provider and data exfiltration succeeding on all three. We propose mitigations for each failure but do not evaluate them. In every failure we observed, the explanation was fluent and wrong. Moreover, nothing in the explainer architecture checks whether an explanation is true before an operator acts on it. Testing the explainer therefore belongs in any audit of an agentic deployment.

著者のコメント

13 pages, 4 figures. Accepted at the ICML 2026 Workshop on Failure Modes in Agentic AI

arXiv ID: 2609.23215 / 要約の誤りについて