arXiv論文メモ
新着一覧
cs.GR / cs.CV / cs.RO / cs.SD · 査読状況未確認

音の方向と文章を使って人体動作を生成するMoSAT

MoSAT: Human Motion Generation from Spatial Audio and Textual Description

Shuyang Xu, Zhiyang Dou, Yiduo Hao, Zekun Li, Liang Pan, Jingbo Wang, Cheng Lin, Yuan Liu, Wenping Wang, Mingmin Zhao, Taku Komura

この論文をやさしく読む

ひとことで言うと

文章で行動を指定するだけでなく、どの方向から音がするかも取り入れて、人の全身動作を生成する方法です。

何に役立つ?

考えられる用途は、音のする方向に反応する人物アニメーションなどの生成です。動作・音響・文章を揃えたデータセットと評価の仕組みも提供します。

この研究の面白いところ

音を単なるタイミングの手掛かりにせず、方向のある環境情報として文章の意図と階層的に組み合わせます。

どこまで分かった?

要旨は最先端の性能を報告しますが、具体的な指標値や比較手法、評価規模は示していません。生成動作の評価を、実際の人やロボットの行動性能と同一視することはできません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

人の動きは、外部の音響事象と行動意図の両方に形作られる。空間音響は反応を引き起こしたり導いたりする環境の手掛かりを伝え、文章は望む行動とその実行の仕方を指定する。本論文では、従来研究でほとんど見過ごされてきた、空間音響と自然言語を同時に条件とする人体動作合成という新しい課題を研究する。 この課題を支えるため、動作系列に空間音響と詳細な文章注釈を対応付けたデータセットSTAMを導入する。その豊かな語彙により、人の動きを正確かつ細やかに指定できる。さらに、自然言語の意図と方向を持つ空間音響の手掛かりを、動作生成前の階層的な交差注意を通じて共同条件とする、全身動作生成用の潜在フローマッチングの枠組みMoSATを導入する。この階層設計は、時間的な一貫性と意味的な対応を備えた動作系列を促進する。 また、この新しい課題を包括的に評価するため、3モダリティの評価器も開発する。広範な実験により、MoSATは文章の意味とともに、空間音響が本来持つ動作を形作る性質を活用し、最先端の性能を達成して、さまざまな状況で正確かつ多様な動作を可能にすることを示す。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-20(UTC)
最新改訂
2026-09-20 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Human motion is shaped by both external acoustic events and behavioral intent: spatial audio conveys environmental cues that elicit or guide a response, while text specifies the desired action and how it should be performed. In this paper, we study the novel task of human motion synthesis jointly conditioned on spatial audio and natural language, a problem that has been largely overlooked in previous research. To support this task, We introduce STAM, a dataset of motion sequences paired with spatial audio and detailed textual annotations whose rich vocabulary affords precise and nuanced specification of human motions. We further introduce MoSAT, a latent flow-matching framework for full-body motion generation jointly conditioned on natural-language intent and directional spatial-audio cues through hierarchical cross-attention before generating motion. Such a hierarchical design enhances temporally coherent and semantically aligned motion sequences. We also develop tri-modal evaluators for comprehensive evaluation on this novel task. Extensive experiments show that MoSAT achieves the SOTA performance by leveraging spatial audio's intrinsic motion-shaping properties alongside textual semantics, enabling precise and diverse motion in various scenarios.

arXiv ID: 2609.23797 / 要約の誤りについて