音声会話と同期するリアルタイム対話アバターの公開基盤
AVTR-1: Open Stack for Real-Time Interactive Avatars
この論文をやさしく読む
ひとことで言うと
会話相手の音声にも反応し、映像と音声を同期して配信するリアルタイム対話アバターの基盤を示した研究。
何に役立つ?
音声エージェントとアバターを組み合わせた対話システムの構築や、聞き手の動きが相手の発話を反映しているかの評価に役立つ。
この研究の面白いところ
動作生成だけでなく再生時刻や割り込みまで含む基盤を扱い、相手の音声が聞き手の動作に使われたかを測る指標も提案している。
どこまで分かった?
要旨では比較対象内の品質指標と2種類の商用音声エージェントでの遅延検証を報告している。すべての運用環境での遅延保証は示されていない。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
話す顔の映像を作るモデルや2人の対話を扱うモデルは、リアルタイム推論が可能になっている。しかし、動きを速く生成するだけでは対話型の会話システムにはならない。稼働中のシステムには、外部の音声エージェントの発話とモデル出力の同期、動画フレームの再生時刻の調整、割り込みへの対応が必要である。本研究は、両参加者の音声を条件とする1億5300万パラメータの自己回帰型フローマッチング動作生成器を中心に、リアルタイム対話アバターの公開基盤 AVTR-1 を導入する。音声エンコーダーは自己蒸留によりストリーミング向けに適応させた。 この基盤は、モデルの塊ごとの生成結果を、外部音声エージェントが駆動する連続的で同期した音声・映像ストリームに変換する。利用者が感じる遅延に対する基盤の寄与を解析的に導き、商用音声エージェント2種類でその境界を検証した。追加の実験では、比較した2者対話システムの中で、報告されたすべての映像品質指標と従来の聞き手動作指標の大半で先行し、口の動きと音声の同期でも競争力を保った。データセンター用と一般消費者用のGPUで、推論はリアルタイムで動作する。 ただし従来の聞き手動作指標では、相手の発話が生成動作に寄与したかは分からない。そこで、聞き手の履歴と話し手の動作を考慮した後でも、話し手の発話に追加の予測情報があるかを測る Reference-Based Directed Granger Gain(R-DGG)を導入する。R-DGG は、記録された聞き手と評価対象の全2者対話システムでは統計的に裏付けられた予測依存性を検出したが、対になる音声を持たない話す顔の生成器や、対応を入れ替えた話し手・聞き手の組では検出しなかった。モデルの重み、描画器、配信バックエンドを、それぞれの部品に応じたライセンスで公開する。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-19(UTC)
- 最新改訂
- 2026-09-19 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-19 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Talking-head and dyadic models now achieve real-time inference, yet fast motion generation alone does not produce an interactive conversation. A live system must synchronize the model's output with speech from an external voice agent, schedule video frames for playback, and handle interruptions. We introduce AVTR-1, an open stack for real-time interactive avatar conversations, built around a compact 153M-parameter autoregressive flow-matching motion generator conditioned on both participants' audio. We adapt its audio encoder for streaming through self-distillation. The stack turns the model's chunk-based generation into a continuous, synchronized audio-video stream driven by an external voice agent, and we analytically derive its contribution to the user-facing latencies and validate the resulting bounds with two commercial voice agents. Further experiments demonstrate that AVTR-1 leads the compared dyadic systems on all reported visual-quality metrics and most conventional listening-motion metrics while remaining competitive in lip synchronization. Its inference runtime operates in real time on data-center and consumer GPUs. However, conventional listening metrics do not establish whether the paired speaker's speech contributes to generated motion. We therefore introduce the Reference-Based Directed Granger Gain (R-DGG), which measures the additional predictive information carried by speaker speech after accounting for listener history and speaker motion. R-DGG finds statistically supported predictive dependence for recorded listeners and all evaluated dyadic systems, but not for talking-head generators without paired audio or mismatched speaker-listener pairs. We release the model weights, renderer, and serving backend under component-specific licenses.
arXiv ID: 2609.22913 / 要約の誤りについて