操作に応じて映像を生成するAstronex-World
Astronex-World 1.0: Real-Time Interactive World Model Foundation
この論文をやさしく読む
ひとことで言うと
カメラの動きや連続的な行動、途中の文章イベントに応じて、次の映像を生成する世界モデルです。
何に役立つ?
操作に反応する映像環境を構築する用途が考えられます。身体を持つAIや自動運転への追加学習向けの入出力も用意しています。
この研究の面白いところ
全体を見て生成するモデルと、過去の計算を再利用して連続生成するモデルを組み合わせた系列です。5Bモデルを段階的に学習し、1台のL20で832×480、24fpsのリアルタイム生成を報告しています。
どこまで分かった?
比較優位はWBenchのNaviとFullなどの記載された評価での結果です。行動インターフェースの用意は、自動運転や身体制御の安全性・成功を実証したこととは別です。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
操作可能な動画世界モデルのオープンな基盤、Astronex-World 1.0を提示する。テキスト指示からの動画生成、または初期観測画像からの動画生成において、フレームに対応したカメラ軌道、連続的なアクション、身体形態の識別子のもとで将来の視覚状態を予測し、展開列の指定位置に挿入するテキストイベントも受け付ける。 モデル群は、文脈全体を用いて生成する双方向モデルと、ブロック因果アテンションおよびブロック間のKVキャッシュを用いて生成を持続する因果モデルを提供する。いずれもWan2.2-TI2V-5Bの事前学習モデルを基に構築する。PRoPEがカメラの内部・外部パラメータを注入し、64次元のアクション系列がすべてのTransformer層を変調する。5段階の学習では、双方向のカメラ・アクション制御を獲得し、バックボーンをブロック因果生成に変換し、少数ステップの生徒モデルへ蒸留し、複数領域のダイナミクスを回復させ、非対称のDMD/DMD2分布整合を適用する。 因果モデルは832×480の動画を24 fpsで生成する。5段階すべての学習はNVIDIA L20 48 GB GPUを2基用いて実行し、因果モデルは1基でリアルタイム配信できる。WBench Naviで73.5、WBench Fullで70.0を記録した。Fullでは、この50億パラメータのモデルは136億のLongCat-Videoと140億のHeliosを上回り、220億のLTX-2.3との差は1点以内だった。また、同じ50億パラメータの事前学習モデルからNVIDIA A100 GPUで追加学習されたYUME 1.5も上回った。あらかじめ設けたアクション入出力インターフェースによって、身体性知能や自動運転に向けた追加学習が可能になる。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-17(UTC)
- 最新改訂
- 2026-09-17 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-17 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
We present Astronex-World 1.0, an open controllable video world-model foundation. Given a text prompt (text-to-video) or an initial observation (image-to-video), the model predicts future visual states under frame-aligned camera trajectories, continuous actions, and an embodiment identifier, and accepts text events inserted at a specified position of a rollout. The family provides a bidirectional model for full-context generation and a causal model with block-causal attention and cross-block KV caching for persistent generation, both built on the Wan2.2-TI2V-5B prior. PRoPE injects camera intrinsics and extrinsics, while a 64-dimensional action stream modulates every Transformer layer. A five-stage training path develops bidirectional camera and action control, converts the backbone to block-causal generation, distills a few-step student, restores mixed-domain dynamics, and applies asymmetric DMD/DMD2 distribution matching. The causal model generates 832x480 video at 24 fps. All five training stages run on two NVIDIA L20 48 GB GPUs, and the causal model streams in real time on one. It scores 73.5 on WBench Navi and 70.0 on WBench Full. On Full, this 5B model is above the 13.6B LongCat-Video and the 14B Helios, within one point of the 22B LTX-2.3, and above YUME 1.5, which is post-trained from the same 5B prior on NVIDIA A100 GPUs. The reserved action input and output interfaces allow post-training for embodied intelligence and autonomous driving.
著者のコメント
Technical report. 25 pages, 13 figures, 10 tables. Project page: https://world.astronex.com.cn ; Code: https://github.com/Astronex-Robotics/Astronex-World ; Weights: https://huggingface.co/Astronex-Lab/Astronex-World
arXiv ID: 2609.20034 / 要約の誤りについて