arXiv論文メモ
新着一覧
cs.CV / cs.AI / cs.HC / cs.LG · 査読状況未確認

3Dアバターの動きを線形近似しモバイルで最大60fpsを実現

One Basis to Animate Them All: Gaussian Blendshape Distillation for Real-Time Avatars

Ramazan Fazylov, Stamatis Lefkimmiatis, Ivan Laptev

この論文をやさしく読む

ひとことで言うと

3Dアバターを毎フレーム重いニューラルモデルで動かす代わりに、少数の基本変形を混ぜて動かす方法です。基本変形の混合係数だけを軽いネットワークで予測します。

何に役立つ?

考えられる用途は、モバイル端末など計算資源の限られた環境でのリアルタイムアバターです。評価では表情と衣服を含む全身動作を扱い、最大60fpsを報告しています。

この研究の面白いところ

個体ごとに別の基本変形を作るだけでなく、個体に依存しない共通の線形構造を利用しています。メモリー予算と描画への影響を考慮して基底を作ります。

どこまで分かった?

検証したのは三つのアバターモデルです。最大3桁の削減と60fpsは最大値であり、要旨には端末名や条件別の値はありません。元モデルの再学習は不要ですが、係数予測用の浅いネットワークは学習します。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

3Dガウシアンアバターは高速な描画に対応するが、計算コストの高いニューラル推論がリアルタイムの動作生成を妨げることが多い。本研究ではこのボトルネックに対処し、事前学習済みアバターモデルの動作が、個体に依存しないブレンドシェイプの線形結合で精密に近似できることを示す。この知見に基づき、フレームごとの重いニューラル復号を浅い係数予測器と線形ブレンドに置き換える蒸留手法、GALA(Gaussian Animation via Linear Approximation)を導入する。 忠実度を高め、必要メモリーを減らすため、描画を考慮した距離尺度とメモリー予算の下で、ブロック局所的な主成分分析を用いて基底を構成することを提案する。本手法はブレンドシェイプ係数を予測する浅いMLPネットワークを学習し、元モデルを再学習することなく、さまざまな動作生成アーキテクチャに適用できる。 表情と、衣服の動きを含む全身の3D動作生成に用いる三つの異なるアバターモデルの推論を高速化して、GALAを検証する。これらのモデルで、蒸留は学習に含めなかった個体にも一般化し、描画品質の大部分を保ちながらCPUの動作生成コストを最大3桁削減する。良好な結果は、学習済みアバター表現が共通の線形構造を持つことを裏付け、モバイル端末で最大60fpsに達する、高効率で正確な動作生成を可能にする。プロジェクトページはhttps://ramazan793.github.io/gala/である。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-10-01(UTC)
最新改訂
2026-10-01 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

3D Gaussian avatars support fast rendering, however, their real-time animation is often challenged by the costly neural inference. We address this bottleneck and show that the animation of pretrained avatar models can be closely approximated by a linear combination of identity-independent blendshapes. Building on this finding, we introduce GALA (Gaussian Animation via Linear Approximation), a distillation method that replaces per-frame heavy neural decoding with a shallow coefficient predictor and a linear blend. To improve fidelity and reduce memory requirements, we propose to construct the basis using block-local PCA under a rendering-aware metric and a memory budget. Our method learns a shallow MLP network to predict blendshape coefficients and applies to various animation architectures without retraining original models. We validate GALA by accelerating the inference of three distinct avatar models for 3D animation of facial expressions and full-bodies with clothing dynamics. Across these models, our distillation generalizes to held-out identities and reduces CPU animation cost by up to three orders of magnitude while preserving most of the rendering quality. Excellent results of our method confirm the shared linear structure of learned avatar representations and enable highly efficient and accurate animation at frame rates reaching up to 60fps on mobile devices. Project page: https://ramazan793.github.io/gala/

arXiv ID: 2610.02207 / 要約の誤りについて