文章から動作を一度の計算で生成するMixiMotion
MixiMotion: One-Step Text-to-Motion Generation via Asymmetric Set Distillation
この論文をやさしく読む
ひとことで言うと
文章から人の動きを作る処理を、繰り返し計算せず一度のネットワーク評価で行う手法です。
何に役立つ?
文章からの動作生成で、生成品質と処理時間の両方を評価する際の候補になります。
この研究の面白いところ
複数の教師動作を集合として学習し、教師に近い品質を維持しながら、遅延を829.58ミリ秒から9.30ミリ秒へ減らしました。
どこまで分かった?
評価はViMoGenと人手評価に基づき、教師のスコアにはまだ届いていません。要旨に実機動作での評価はありません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
文章から動作を繰り返し生成する方法は高品質で意味にもよく合うが、ネットワークを何度も評価するため推論に時間がかかる。著者らは、オフラインの集合蒸留に基づき、厳密に一回の計算で文章から動作を生成するMixiMotionを提案する。各文章の指示に対して教師モデルの軌道を一つだけ蒸留するのではなく、教師が生成した複数の動作をオフラインの集合として保存し、教師と生徒の標本集合を非対称な双方向の対応付けで合わせる。教師から生徒への方向は教師が支持する多様な動作を網羅するよう促し、生徒から教師への方向は教師が支持しない生成を抑える。さらに、復号した動作空間で微分可能な運動学的監督を加え、正規化表現の一致を実際の動作空間の制約で補う。推論時には教師への問い合わせ、反復サンプリング、候補の順位付けをせず、一度のネットワーク評価で完全な動作列を生成する。ViMoGenでは意味的一致のスコアが0.835で、評価した一段階手法を上回り、50段階の教師HY-Motion-1.0-Liteの0.858に近づいた。評価者に方法を伏せた人手評価では総合点が4.33で、教師の4.50には及ばないものの、評価した一段階または少数段階の基準法を上回った。一方、生成の遅延は829.58ミリ秒から9.30ミリ秒に減り、89.2倍の高速化に相当する。これらは厳密な一段階の文章から動作への生成で、品質と効率を両立できることを示す。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-19(UTC)
- 最新改訂
- 2026-09-19 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-19 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Iterative text-to-motion generation delivers high-quality and semantically aligned motions but requires multiple network evaluations, resulting in substantial inference latency. We present \textbf{MixiMotion}, a strict one-step text-to-motion generation framework based on offline set distillation. Instead of distilling a single teacher trajectory for each text prompt, MixiMotion constructs an offline bank of multiple teacher motions and aligns teacher and student sample sets through \textbf{asymmetric bidirectional matching}. The teacher-to-student direction promotes coverage of diverse teacher-supported motions, while the student-to-teacher direction suppresses unsupported generations. We further introduce differentiable decoded-space kinematic supervision to complement normalized representation matching with constraints in the decoded motion space. At inference, MixiMotion generates a complete motion sequence with a single network evaluation, without teacher queries, iterative sampling, or candidate ranking. On ViMoGen, MixiMotion achieves a semantic alignment score of $0.835$, outperforming the evaluated one-step baselines and approaching the $0.858$ score of its 50-step HY-Motion-1.0-Lite teacher. In blinded human evaluation, MixiMotion obtains an overall rating of $4.33$, compared with $4.50$ for the teacher, while outperforming the evaluated one-/few-step baselines. Meanwhile, generation latency is reduced from $829.58$\,ms to $9.30$\,ms, corresponding to an $89.2\times$ speedup. These results demonstrate an effective quality--efficiency trade-off for strict one-step text-to-motion generation.
著者のコメント
Under Review
arXiv ID: 2609.23010 / 要約の誤りについて