身体性とエージェント協調を統合する視覚言語モデル
ME-VLM: A Unified VLM for Embodied Cognition and Agent Coordination
この論文をやさしく読む
ひとことで言うと
物理環境での知覚・計画とマルチモーダルエージェントの能力を、一つの視覚言語モデルへ統合した研究。
何に役立つ?
身体的な課題とデジタルのエージェント課題を同じモデルで扱う設計や、端末上での推論最適化の参考になる。
この研究の面白いところ
4B版をM100上で動かし、事前入力処理の遅延を400ミリ秒から188ミリ秒へ短縮した。
どこまで分かった?
要旨では各ベンチマークの具体的な精度値は示していない。遅延の数値はM100上の4B版についての結果である。
v2のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
物理世界で働くAIには、環境の制約と実行結果からのフィードバックを踏まえて、視覚と言語の理解を実際の環境に結びつけることが必要である。本研究は、身体的な認知とマルチモーダルなエージェントの能力を統合する視覚言語モデルMachEmbodied-VLM(ME-VLM)を提案する。モデルには4Bと35B-A3Bの二種類がある。物理的な知覚と時空間的な推論に加え、デジタル環境と物理環境の両方での計画、対話、結果の評価を重視する。 学習データは身体的な課題とマルチモーダルエージェントの課題を含み、結果の評価や判断の修正を支える実行時の観測とフィードバックも組み込む。学習過程は、身体的能力の注入、身体的能力とマルチモーダルエージェント能力それぞれの専門モデルの強化学習、そして相補的な能力を単一モデルにまとめる複数教師のオンポリシー蒸留から成る。実験では、身体的課題とエージェントのベンチマークに加え、自動運転と身体的ナビゲーションの課題でも競争力のある性能を示した。端末での実行に向けて、視覚トークンの圧縮、W4A8量子化、ハードウェアとソフトウェアの共同最適化により、4B版をM100上で推論可能にし、事前入力処理の遅延を400ミリ秒から188ミリ秒へ短縮した。プロジェクトページ:https://machembodied.com/ME-Brain/ME-VLM.html。コード:https://github.com/MachEmbodied/ME-VLM。
v2の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-21(UTC)
- 最新改訂
- 2026-09-22 · v2
- 査読・掲載
- 査読状況未確認
更新履歴
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Physical AI requires models to ground visual and linguistic understanding in real-world environments while accounting for environmental constraints and execution feedback. We introduce MachEmbodied-VLM (ME-VLM), a unified vision-language model with two variants, 4B and 35B-A3B, that brings together embodied cognition and multimodal agent capabilities. Our work emphasizes physical perception and spatiotemporal reasoning, together with planning, interaction, and outcome assessment in both digital and physical environments. We construct training data spanning embodied and multimodal agent tasks, including execution observations and feedback to support outcome assessment and decision refinement. The training pipeline comprises embodied capability injection, separate reinforcement learning of embodied and multimodal-agent experts, and multi-teacher on-policy distillation that consolidates their complementary capabilities into a single model. Experiments show competitive performance on both embodied and agent benchmarks, as well as on autonomous-driving and embodied-navigation tasks. For edge deployment, visual token compression, W4A8 quantization, and hardware-software co-optimization enable on-device inference of the 4B variant on the M100, reducing prefill latency from 400 ms to 188 ms. Project Page: https://machembodied.com/ME-Brain/ME-VLM.html Code Repository: https://github.com/MachEmbodied/ME-VLM
arXiv ID: 2609.24526 / 要約の誤りについて