arXiv論文メモ
新着一覧
cs.AI / cs.CL · 査読状況未確認

道案内に失敗するAIにも内部地図はあるのか

World Modeling in Transformers

Pierre Beckmann, Matthieu Queloz, Andre Freitas

この論文をやさしく読む

ひとことで言うと

AIが道案内に失敗しても、内部に地図を持っていないとは限らない、という問題を調べています。TaxiGPTでは、地図に相当する表現はある一方、交差点の特徴の干渉で位置を取り違えると報告しています。

何に役立つ?

AIの能力を成功率だけで判断せず、環境の表現、位置追跡、行動への利用を分けて評価するための手掛かりになります。失敗原因を内部機構から調べる研究です。

この研究の面白いところ

内部の表現を観察するだけでなく、因果的な介入で役割を調べています。同じ方向へ動ける交差点をまとめる表現が、位置の誤りの影響を抑えるという点も特徴です。

どこまで分かった?

中心となる分析対象はマンハッタンのランダムウォークで学習したTaxiGPTです。この結果から、すべてのTransformerが現実世界について忠実な世界モデルを持つとはいえません。要旨には各指標の具体的な値はありません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

行動上の失敗によって、Transformerが環境を忠実に表す表現を学習していても、世界モデルを持たないように見えることがある。マンハッタン内のランダムウォークで学習したTransformerであるTaxiGPTを用いて、これを示す。このモデルの失敗は、内部地図に一貫性がない証拠として解釈されてきた。 内部機構の分析と因果的な介入を通じて、このモデルが交差点と道路を表現し、自分の位置を追跡し、目標方向を示すコンパスを使って経路を選ぶことを示す。失敗の原因を追跡すると、重ね合わせて表現された交差点の特徴同士の干渉が、内部地図上の自己位置の特定を乱していることが分かる。同じ合法な移動を持つ交差点の表現をまとめる「アフォーダンス・パッキング」は、こうした誤りの影響を抑えるのに役立つ。 最後に、モデル同士を比較するための内部機構に基づく指標を提案し、世界をモデル化する能力が学習の異なる段階で現れることを示す。この結果は、モデルが世界モデルを持つかどうかを問うことから、モデルが環境を表現し、その表現で行動を導くための相互作用する能力、すなわち世界のモデル化を、内部機構に即して研究することへの転換を促す。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-18(UTC)
最新改訂
2026-09-18 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Behavioral failures can make a transformer appear to lack a world model even when it has learned faithful representations of its environment. We demonstrate this in TaxiGPT, a transformer trained on random walks through Manhattan whose failures have been interpreted as evidence of an incoherent internal map. Through mechanistic analysis and causal interventions, we show that the model represents intersections and streets, tracks its position, and uses a goal compass to navigate. We trace its failures to interference between superposed intersection features, which disrupts localization within the internal map. Affordance packing, which groups representations of intersections with the same legal moves, helps limit the consequences of these errors. Finally, we propose mechanistic indicators that we use to compare models and show that world-modeling capacities emerge at different stages of training. Our findings motivate a shift from asking whether a model has a world model to mechanistically studying its world modeling: the interacting capacities through which it represents its environment and uses those representations to guide behavior.

arXiv ID: 2609.21748 / 要約の誤りについて