arXiv論文メモ
新着一覧
cs.AI / cs.LG · 査読状況未確認

判断モデルに予測を任せずシミュレーション結果を渡す

Code Owns the Simulation, Jev Owns the Evaluation

Yaodong Yang, Hongyao Tang, Yi Ma, Xingyu Fan, Weixun Wang, Jinpeng Li, and Tianpei Yang

この論文をやさしく読む

ひとことで言うと

選択肢を評価するAIに、行動の先の展開まで一度に考えさせると失敗しやすく、別のコードで予測結果を渡すと改善するという研究です。

何に役立つ?

考えられる用途は、ゲームやロボットの行動選択で、シミュレーターと判断モデルの役割を分ける設計です。評価対象では、先読みや物理シミュレーションを外部で行って結果を与える方法が有効でした。

この研究の面白いところ

失敗の原因が単なる知識不足ではない点です。同じモデルでも、相手の行動を単独で質問すれば答えられるのに、予測と選択を1回の呼び出しにまとめると誤るという違いを示しています。

どこまで分かった?

99%は認知的熟慮テストに関する値で、すべての課題での成功率ではありません。ロボット制御を含む各条件の具体的な成績や実機検証の範囲は要旨には示されていません。原要旨のモデル名は未展開の記号で、訳ではタイトルのJevを用いています。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

Jevのような判断モデルは、推論文を出力せず、1回の呼び出しで記述された各選択肢の確率を返す。この性質は、エージェントの行動選択層として魅力的だが、どの判断を信頼して任せられるかは明らかでない。本研究ではJevを、熟慮を問うテスト、1回限りの行列ゲーム、テキストゲームALFWorld、ロボット制御で評価し、明確な境界を見いだす。入力に書かれた内容から正しい選択肢を判断できる場合、Jevは成功する。本研究ではこれを「評価」と呼ぶ。具体的には、直感に反する答えを持つ認知的熟慮テストの問題の99%を解く。 一方、正しい選択肢が「シミュレーション」、すなわち入力にない事柄を予測することに依存すると失敗する。たとえば、相手の行動や、先に達成すべき下位目標の予測が該当する。ゲームでは相手の行動が与えられていないため、Jevは合理的な相手がランダムに行動するかのように、最適でない手を選ぶ。ALFWorldでは、課題文にある物や場所に言及する命令を好む。たとえば「きれいなナイフを引き出しに入れる」という課題に対し、先に流しで洗う代わりに、洗っていないナイフを直接引き出しへ運ぶ。 意外にも、これらの失敗の多くは知識不足によるものではない。相手が何をするかを別に尋ねると、Jevはたいてい正しく答え、相手の行動が与えられれば適切に応答する。失敗するのは、1回の呼び出しでシミュレーションを行い、さらにその結果に基づいて評価しなければならない場合である。このことは、予測やシミュレーションをコードに任せるべきだと示唆する。ALFWorldの先読みやロボット制御の物理シミュレーションのようにコードが結果を供給すると、Jevは汎用的な評価能力によって熟練した制御器となる。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-10-01(UTC)
最新改訂
2026-10-01 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Judgment models such as \jev{} return, in a single call and without reasoning text, a probability for each described option. This makes them attractive as an agent's action-selection layer, but it is unclear which decisions they can be trusted with. We test \jev{} on reflection tests, one-shot matrix games, the text game ALFWorld and robot control, and find a sharp boundary. \jev{} succeeds when the right option can be judged from what the input describes, which we call \emph{evaluation}. Specifically, it solves 99\% of the counterintuitive Cognitive Reflection Test questions. However, it fails when the right option depends on \emph{simulation} (i.e., predicting something not in the input), such as the opponent's action or the subgoal that must come first. In games, \jev{} plays suboptimally as if its rational opponent acted at random, because the opponent's action is not given. In ALFWorld, \jev{} favors commands that mention an object or place named in the task description. For example, given the task ``put a clean knife in the drawer'', \jev{} carries an unwashed knife straight to the drawer instead of first washing it at the sink. Surprisingly, many of these failures are not due to a lack of knowledge. Asked separately what the opponent will do, \jev{} usually answers correctly, and it responds well given the opponent's action. It fails when one call must both perform the simulation and evaluate based on it. This suggests letting code make the prediction or simulation. When code supplies it, such as a lookahead in ALFWorld and physics simulation in robot control, \jev{} becomes an expert controller through its general evaluation ability.

著者のコメント

10 pages main text, 20 pages total with appendix; 6 figures, 7 tables. Preprint

arXiv ID: 2610.01834 / 要約の誤りについて