会話の状況に合わせて聞く・話す時間を変える音声対話モデル
AdaptDuplex: from static to adaptive full-duplex spoken dialogue
この論文をやさしく読む
ひとことで言うと
会話中の状況に応じて、聞く時間と話すタイミングを変える音声対話モデルです。
何に役立つ?
人と話しながら応答する音声システムの、発話の順番や重なりの制御を検討するのに役立ちます。
この研究の面白いところ
各行動判断をトークンにし、実行時の制御と時間窓の動的な選択を両立させています。
どこまで分かった?
要旨の比較は指定された3種類のベンチマークによるものです。実環境での利用全般については示していません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
全二重の音声対話では、刻々と変わる会話のタイミングや認知的な負荷に対応しつつ、1秒未満の遅延で聞くことと話すことを同時に行う必要がある。しかし現行モデルの多くは動作設定が固定され、適応的に判断する体系的な仕組みを欠く。本研究ではQwen3-Omniを拡張し、3層にわたって設計した適応機構を備えるAdaptDuplexを提案する。簡潔なトークン単位のプロトコルで各時間窓を標準的な系列として表し、音声よりテキストを一定範囲で先行させて二つの流れを整合させる。また、行動上の判断を明示的なトークンとして表し、ロジットのバイアスによって追加学習なしで実行時に制御できる。適応機構は離散的な時間窓の長さを動的に予測し、必要に応じて、直接応答に加えて、処理を妨げない認知的な情報統合と複数の外部推論を行う。段階的な処理系では、Thinkerの3段階の学習課程、Talkerのみの教師あり微調整、両者の共同微調整を経て、GRPOを追加する。Full-Duplex-Bench v1とv1.5では、順番交代、発話の重なり、タイミングについて比較可能な指標の大半でDuplexOmniとMiniCPM-o 4.5を上回り、応答するかどうかの判断と応答時刻の双方が改善した。人が録音したHumDial-FDBenchでは、比較した全二重モデル中で最高のFinalスコア72.9を得た。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-24(UTC)
- 最新改訂
- 2026-09-24 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-24 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Full-duplex spoken dialogue requires simultaneous listening and speaking at sub-second latency, under conversational timing and cognitive demands that change moment to moment. Yet current models mostly impose static operating points, lacking a systematic mechanism for adaptive decisions. We present AdaptDuplex, which upgrades Qwen3-Omni with such a mechanism, co-designed across three layers. A compact token-level protocol represents every window as a canonical sequence, trains dual-stream alignment through a bounded text lead over speech, and exposes every behavioral decision as an explicit token for training-free runtime control via logits bias. Adaptive mechanisms dynamically predict among discrete window durations and augment direct response as needed with non-blocking cognitive consolidation and multi-flight external reasoning. A progressive pipeline introduces these behaviors through a three-stage Thinker curriculum, then Talker-only and joint SFT, with GRPO as a further increment. On Full-Duplex-Bench v1 and v1.5, AdaptDuplex outperforms DuplexOmni and MiniCPM-o 4.5 on the majority of comparable turn-taking, overlap-behavior, and timing metrics, with gains in both interaction decisions and response timing. On the human-recorded HumDial-FDBench, it attains the highest Final score (72.9) of the compared duplex models.
arXiv ID: 2609.29217 / 要約の誤りについて