arXiv論文メモ
新着一覧
cs.CV · 査読状況未確認

文章による動作編集で元の動きの細部を残すSuperMotion

SuperMotion: Source-Preserving Denoising for Text-Driven Human Motion Editing

Fa-Ting Hong, Peter Wonka

この論文をやさしく読む

ひとことで言うと

文章で人の動作を編集するとき、変更が必要な部分以外の細かい動きを、元の動作から毎段階取り込んで残す方法です。

何に役立つ?

考えられる用途は、動作データの一部を文章で指定して編集する作業です。要旨ではMotionFixなどで、要求された変更の精度と時間的な細部の保持を評価しています。

この研究の面白いところ

元の動きを条件として入力するだけでなく、ノイズ除去の各段階で出力に直接混ぜます。どこをどれだけ再利用するかを、明示的な編集マスクなしで学習します。

どこまで分かった?

報告された33.20%はMotionFixの候補プール全体に対するR@1という評価値で、任意の編集要求の成功率を表すものではありません。要旨には処理速度や実利用者による評価は記載されていません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

文章に基づく人の動作編集は、元の内容のうち変更と両立する部分を保ちながら、要求された変更を実現することを目指す。既存の拡散型編集手法は、編集しない部分の保持を主として学習された条件付けに頼っているが、ノイズ除去が進むにつれて、出力が時間的な細部を失うことがある。本研究は、元データを保持するため、逆過程の各ステップで元データを明示的に再利用するSource-Preserving Denoising framework(SuperMotion)を提案する。 まず元の動作を出力の時間軸に合わせ、フレームと特徴次元にわたる再利用を制御する保持ゲートを予測する。次に、ノイズのない空間での元データのアンカーが、学習した保持ゲートを使い、予測されたノイズのない動作と時間位置を合わせた元の動作を混合し、補正後の推定値をサンプリングの事後分布へ直接渡す。時間位置を合わせた元データは回帰出力ではなく、実現された動作であるため、このアンカーは、再構成で訓練したノイズ除去器が平滑化して失いやすい、個々の標本に備わる時間的な細部を注入する。 有効な再利用を学ぶため、アンカーで補正した推定値を編集対象の正解に対して教師ありで学習し、時間方向の高周波損失を使って2階時間差分も一致させる。これらの目的関数は明示的な編集マスクを必要としない。広範な実験は、SuperMotionが編集精度を改善し、MotionFixで候補プール全体に対するR@1が33.20%に達することを示す。同時に、時間的な細部の減衰を抑え、要求された変更を実現しながら動作の動的特性を保つ。アブレーション分析により、改善をもたらすのは学習した保持ゲートであり、それが元データを再利用して、編集しない内容を適切に保持することを確認する。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-10-01(UTC)
最新改訂
2026-10-01 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Text-driven human motion editing aims to realize a requested change while preserving compatible source content. Existing diffusion editors rely largely on learned conditioning for preservation of the unedited part, yet their outputs can lose temporal detail as denoising proceeds. We propose the \textbf{Source-Preserving Denoising framework (SuperMotion)}, which explicitly reuses the source at each reverse step for source preservation. We first align the source motion to the output timeline and predict a preservation gate that controls reuse across frames and feature dimensions. A clean-space source anchor then utilizes the learned preservation gate to blend the predicted clean motion with the aligned source and passes the corrected estimate directly to the sampling posterior. Because the aligned source is a realized motion rather than a regression output, the anchor injects sample-level temporal detail that a reconstruction-trained denoiser tends to smooth away. To learn effective source reuse, we supervise the anchored estimate against the editing target and match its second temporal differences through a temporal high-frequency loss. These objectives require no explicit edit masks. Extensive experiments show that SuperMotion improves editing accuracy, reaching 33.20\% full-pool R@1 on MotionFix, while reducing temporal-detail attenuation and preserving motion dynamics as it realizes the requested changes. Ablations confirm that the learned preservation gate is responsible for the gain and that it reuses the source to retain the unedited content properly.

著者のコメント

Under review

arXiv ID: 2610.01517 / 要約の誤りについて