arXiv論文メモ
新着一覧
cs.CV · 査読状況未確認

人の動作予測に向けた潜在空間でのデータセット蒸留

Latent Dataset Distillation for Human Motion Prediction

Ge Tian, Guang Li, Takahiro Ogawa, Miki Haseyama

この論文をやさしく読む

ひとことで言うと

人の動作を予測する学習用データを、自然な動作の形を保ちながら小さな合成データへ圧縮する方法を提案しています。

何に役立つ?

考えられる用途は、動作予測モデルの訓練に使うデータを少数にすることです。要旨で示された効果は三つのデータセットと二種類の予測モデルでの比較結果です。

この研究の面白いところ

動作を直接合成する代わりに、学習済みのデコーダーが出力できる潜在空間で蒸留し、残差量子化で表現力を高めています。

どこまで分かった?

直接の勾配マッチングに対する優位は30条件中27条件であり、全条件ではありません。自然さの比較は要旨では定性的な観察として報告されています。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

データセット蒸留は、下流の学習に役立つ性質を保ちながら、大きな訓練データ集合を小さな合成データ集合へ圧縮する。画像では広く研究され、最近では時系列予測にも拡張されたが、人の動作予測への応用はほとんど調べられていない。人の動作は高次元で、各要素が構造的に結び付いている。元の動作空間での勾配マッチングは、姿勢のもっともらしさや時間的な動きに関する事前知識なしに、多くの相関した変数を最適化するため、現実味に欠け不安定な合成動作を生じやすい。この問題に対し、学習した動作の事前分布で蒸留を制約する潜在空間のデータセット蒸留の枠組みを提案する。まず残差量子化変分オートエンコーダー(RVQ-VAE)で動作を圧縮し、次に固定した量子化器とデコーダーを通じて、学習可能な潜在表現の集合だけを蒸留時に更新する。事前学習済みのデコーダーは合成動作をその出力空間に制限し、残差量子化は複数のコードブックを通じて潜在表現の近似を段階的に改善して、単段のベクトル量子化による表現上の制約を緩和する。Human3.6M、CMU、3DPWのデータと二つの予測モデルを用いた実験では、評価した30条件中27条件で直接の勾配マッチングを上回り、全条件でランダムに選んだ部分集合を上回った。定性的な比較でも、より自然に見える合成動作が得られた。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-22(UTC)
最新改訂
2026-09-22 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Dataset distillation (DD) compresses a large training set into a compact synthetic set while preserving downstream training utility. Although DD has been widely studied for images and recently extended to time-series forecasting, its application to human motion prediction remains largely unexplored. Human motion is high-dimensional and structurally coupled, and gradient matching (GM) in the original motion space optimizes many correlated variables without a prior on pose plausibility or temporal dynamics, which frequently yields implausible and unstable synthetic motions. To address this limitation, we propose a latent DD framework that regularizes distillation with a learned motion prior. Motions are first compressed by a residual-quantized variational autoencoder (RVQ-VAE), and distillation then updates only a learnable latent bank through the frozen quantizer and decoder. The pretrained decoder restricts synthetic motions to its output space, while residual quantization progressively refines the latent approximation across multiple codebooks and alleviates the representational bottleneck of single-stage vector quantization. Experiments on Human3.6M, CMU, and 3DPW with two prediction backbones show that the proposed framework outperforms direct GM in 27 of 30 evaluated settings and random subsets in every setting, and produces visibly more plausible synthetic motions in qualitative comparisons.

arXiv ID: 2609.26430 / 要約の誤りについて