固定した世界モデルで遠い目標へ進む中間目標の選び方
Aim Short to Reach Far: Your Frozen World Model Can Plan Better Than You Think
この論文をやさしく読む
ひとことで言うと
画像から行動を計画するAIで、最終地点の前に中間目標を置くと遠い目標へ進みやすくなると示した。
何に役立つ?
学習済みの世界モデルを再学習せずに、長い行動計画の性能を改善する方法を検討できる。
この研究の面白いところ
予測モデルそのものではなく、行動を採点する目標画像の選択だけで結果が変わった。
どこまで分かった?
改善は要旨に記載されたLeWMモデルと四つの課題での評価。中間目標の距離や検索区間の調整にも依存する。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
視覚的な世界モデルを使う計画器は、予測した各結果を、符号化した最終目標画像との距離で評価することが多い。本研究は、この評価対象が、力学モデルが正確で短期の探索が全域最適であっても制御を制限しうることを示す。目標へ到達するには、最初に目標から遠ざかる行動が必要な場合があるためである。学習済みのLeWMモデルを固定した実験では、中間目標を使うと、Cube、PushT、Reacher、TwoRoomで行動の生成と記録済み行動の順位付けが大きく改善した。学習した中間目標でも、観測済みの経験から選んだ目標でも改善が得られた。 提案するAnchored Planningは、始点と終点の観測が現在と最終目標に似た記録済みの行動区間を検索し、その始点の少し後の観測を中間目標にする。固定したモデルは、現在の状態からその目標へ向かう行動を評価する。追加学習なしで、観測済みの中間目標を使った計画は、長距離の評価に含まれるすべての課題で公開済みのLeWM計画器を上回った。最終目標だけへの探索を増やしても同じ改善は得られなかった。次状態の予測誤差が小さいことは、必ずしも制御の改善につながらない。成功は、中間目標をどれほど先に置くか、実行が進むにつれて検索する区間を短くするかにも依存する。同じ固定モデルと計画器でも、目標の置き方だけを変えることで、最終目標の評価では到達できない目標に届いた。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-24(UTC)
- 最新改訂
- 2026-09-24 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-24 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Planners built on visual world models commonly score each predicted outcome by its distance to the encoded goal image. We show that this target can limit control even with exact dynamics and globally optimal short-horizon search: reaching a goal may require actions that initially move away from it. With frozen LeWM models, intermediate targets substantially improve action synthesis and recorded-action ranking on Cube, PushT, Reacher, and TwoRoom. Learned targets and targets drawn from observed experience both produce these gains. We introduce Anchored Planning, which retrieves a recorded segment whose start and end resemble the current and goal observations, then aims at an observation shortly after its start. The frozen model scores actions toward this target from the current state. Without additional training, planning toward observed targets outperforms the released LeWM planner on every task in our long-range evaluation. Additional final-goal search falls short of the same gains. Lower successor-prediction error need not translate into better control. Success also depends on how far ahead the target is placed and on shrinking the retrieval span as execution advances. Changing only the target lets the same frozen model and planner reach goals that final-goal scoring misses.
arXiv ID: 2609.30036 / 要約の誤りについて