arXiv論文メモ
新着一覧
cs.AI · 査読状況未確認

画像・音声・文章から物理的な状態を理解する小型モデル

OmniFysics-Nano-V2 Technical Report: Understanding the Physical World Across Modalities

Yizhou Liu, Jinghang Han, Kaixiang Qiu, Qi He, Minghao Han, Yue Jiang, Xujia Chen, Wei Zou, Shunli Wang, Lihua Zhang, Dingkang Yang

この論文をやさしく読む

ひとことで言うと

画像や音声などを一緒に扱う小型モデルに、物体の属性や出来事の物理的な関係を学ばせた研究。

何に役立つ?

複数種類の入力から物理状態を理解するモデルを作る際、教師データと学習段階の設計の参考になる。

この研究の面白いところ

物理属性と音響事象を結び付けるデータ作成を加え、比較した21のベンチマークのうち17で最高の結果を報告した。

どこまで分かった?

次世代物理AIの基盤という見通しは著者らの想定であり、要旨にロボットでの実運用結果は示されていない。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

あらゆる種類の入力を扱うモデルによって、画像、音、音声、言語をまたぐやり取りが広がった。しかし学習は意味の説明や汎用的な目的を中心に組まれることが多く、物理的な属性、相互作用の状態、因果的な仕組みは一部しか明示されない。この不足は入力形式を増やすだけでは解消しない。観測を世界の物理的な構造と結び付ける教師信号が必要だからである。本研究は、物理世界の知覚と理解のための小型の全形式対応モデルOmniFysics-Nano-V2を示す。画像、動画、音、音声、文章を共通の推論枠組みで入力でき、文章と音声を生成する。明示的な物理教師信号の不足に対しては、重要な物体を構造化した物理属性に結び付け、画像上の変化を音響事象、中間的な反応、相互作用の結果と対応させる二枝の物理重視データ作成手順を構築する。学習目的の単調さに対しては、報酬の多様性で強化学習用の指示を選び、一般的な課題の正しさから細かな物理知覚の推論へ進む二段階の群相対方策最適化を採用する。複数形式、音と映像、物理推論のベンチマークでは、提案したデータと学習方法により、幅広い全形式対応能力を保ちながら物理世界の理解が改善した。比較した最良水準の全形式対応モデルに対し、21ベンチマーク中17で最高の結果を達成した。著者らは、このモデルが全形式対応の知覚と物理世界の知覚を備え、次世代の物理AIの基盤となる可能性を述べる。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-22(UTC)
最新改訂
2026-09-22 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Omni-modal models have expanded multimodal interaction across vision, audio, speech, and language. However, their training is predominantly organized around semantic descriptions and general-purpose objectives, leaving physical attributes, interaction states, and causal mechanisms only partially specified. This gap is not simply a matter of modality coverage: adding more modalities does not by itself provide the supervision needed to connect observations with the physical structure of the world. We present OmniFysics-Nano-V2, a compact omni-modal model for physical-world perception and understanding. The model supports image, video, audio, speech, and text inputs within a shared reasoning framework, together with text and speech generation. To address the lack of explicit physical supervision, we construct a dual-branch physics-aware data pipeline that grounds salient objects in structured physical attributes and aligns visual changes with acoustic events, intermediate responses, and interaction outcomes. To address homogeneous training objectives, we curate reinforcement-learning prompts by reward diversity and adopt a two-stage Group Relative Policy Optimization curriculum that progresses from general task correctness to fine-grained physical perceptual reasoning. Experiments across multimodal, audio-visual, and physical reasoning benchmarks show that the proposed data and training strategy improves physical-world understanding while preserving broad omni-modal competence. The proposed model achieves leading result on 17 of 21 benchmarks against SOTA omni-modal models. By equipping AI systems with both omni-modal and physical-world perception capabilities, OmniFysics-Nano-V2 is poised to become a cornerstone of next-generation Physical AI.

著者のコメント

24 pages

arXiv ID: 2609.25738 / 要約の誤りについて