arXiv論文メモ
新着一覧
cs.CV / cs.LG · 査読状況未確認

予測の弱点と行動の因果効果を学び直す世界モデル

OnlineWM: Causality-Aware Active Online Learning for Effective World Modeling

Yikun Miao, Fangqi Zhu, Quanxin Shou, Xiaoyi Pang, Zhengyang Yan, Junhao Li, Haodong Wang, Zicong Hong, Song Guo

この論文をやさしく読む

ひとことで言うと

予測を間違えやすい場面をシミュレータから追加収集し、同じ状態で行動だけ変えた結果を比べて世界モデルを改善します。

何に役立つ?

考えられる用途は、行動の違いによる将来予測を安定させることです。要旨ではシミュレータとの相互作用を用いた学習を提示しています。

この研究の面白いところ

データを集め直す対象と、行動が結果を変えたかを学ぶ目的の両方を更新する閉ループ設計です。

どこまで分かった?

要旨には具体的な実験規模や改善率はありません。反実仮想の訓練戦略を採用したことから、任意の実世界環境の因果構造が確定したとはいえません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

生成的な世界モデルは、行動を条件として将来の状態を予測することを目指しており、信頼できる動力学モデルには、行動に応じて予測を制御できることが不可欠である。近年はこの能力を強めるためにシミュレータ生成データを利用する取り組みがあるが、既存の訓練処理には2つの根本的な限界がある。第1に、固定的なオフラインのデータ収集では、訓練データと変化し続けるモデルの誤り方との間に分布の不一致が生じ、動力学予測が不確かな、重要だが稀な場面を解決できない。第2に、観測との食い違いを最小化する標準的な目的は、行動と効果の因果関係を捉える代わりに、見かけ上の相関を利用するようモデルを促すことが多い。 これらに対処するため、シミュレータとの能動的な相互作用と因果関係を考慮した最適化を通じ、世界モデルを継続的に改善するオンライン訓練の枠組みOnlineWMを提案する。OnlineWMは2つの主要な工夫を導入する。1つ目は能動的オンライン学習である。固定データセットを使う代わりに、モデルの現在の予測上の弱点を狙う新しい相互作用系列を、シミュレータに適応的に問い合わせ、有用性の高いデータを取得する。2つ目は因果関係を考慮した微調整である。同一の状態から異なる行動を取った結果を対比する反実仮想学習を提案し、周囲の環境の自然な変化ではなく、特定の行動に状態遷移を帰属させるようにする。これにより、予測を信頼できる因果機構に基づかせる。 能動的なデータ取得と因果的な最適化を統合することで、OnlineWMは閉ループの改良過程を構成し、多様な場面への頑健性と、原因の帰属の精密さを両立させる。広範な実験は、OnlineWMが行動に応じた制御可能性を大きく高め、未見の領域にも有効に汎化することを示す。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-20(UTC)
最新改訂
2026-09-20 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Generative world models aim to predict future states conditioned on actions, where action controllability is fundamental for reliable dynamics modeling. While recent efforts leverage simulator-generated data to enhance this capability, existing training pipelines face two fundamental limitations. First, static offline data collection leads to a distribution misalignment between training sets and the model's evolving error patterns, failing to resolve critical long-tail scenarios where dynamics predictions remain unreliable. Second, the standard objective of minimizing observational discrepancy often encourages the model to exploit spurious correlations instead of capturing the underlying action-effect causality. To address these limitations, we propose OnlineWM, an online training framework that continuously improves world modeling through active simulator interaction and causality-aware optimization. OnlineWM introduces two key innovations: (1) Active Online Learning: Instead of using fixed datasets, OnlineWM adaptively queries the simulator for new interaction sequences that target the model's current predictive weaknesses, ensuring high-utility data acquisition. (2) Causality-Aware Fine-Tuning: We propose a counterfactual learning strategy that contrasts the outcomes of different actions from identical states, forcing the model to attribute state transitions to specific actions rather than ambient environmental evolution, thereby grounding its predictions in reliable causal mechanisms. By integrating active data acquisition with causal optimization, OnlineWM establishes a closed-loop refinement process that ensures the model is both robust to diverse scenarios and precise in its causal attribution. Extensive experiments demonstrate that OnlineWM significantly enhances action controllability and generalizes effectively to unseen domains.

著者のコメント

23 pages, 9 figures. Project page: https://onlinewm.github.io/

arXiv ID: 2609.23753 / 要約の誤りについて