arXiv論文メモ
新着一覧
cs.CL / cs.AI / cs.HC / cs.SD · 査読状況未確認

音声対話の発話交代を話者の意図に応じて評価

Neither Silence nor Overlap Is Failure: Intent-Conditioned Evaluation of Turn-Taking in Full-Duplex Spoken Dialogue Models

Kian Shamsaie, Iman Modarressi

この論文をやさしく読む

ひとことで言うと

会話でいつ話し始めるべきかを、相手の意図も考慮して評価する指標です。

何に役立つ?

同時に聞き話す音声対話モデルの発話交代を、人間の判断に近い形で比較する評価に役立つ。

この研究の面白いところ

沈黙や発話の重なりを一律に失敗とせず、話者の意図に応じて望ましいタイミングを変える。

どこまで分かった?

要旨の比較は五つの対話コーパスと11システムによる。話者意図は注釈に基づく推定であり、直接観測されたものではない。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

全二重音声対話モデルの既存ベンチマークは、直前の発話が完了したかどうかに応じて、すぐ応答するか黙るかを固定時間幅の二値規則で採点する。著者らは、応答の開始を遅らせる沈黙も、相手の発話に先んじて重なる応答も、適切さは話者の潜在的な意図に依存し、その意図は当該話者の行動からしか識別できないと論じる。そこで、二者対話のコーパス5つから9,728の場面、計73.2時間を集めたベンチマークTACTを導入する。各場面には対話履歴、話者ごとの記憶プロファイル、注釈者から推定した6種類の意図についての事後確率を付ける。採点には二値の時間窓に代えて、閾値で重み付けした厳密に適正な連続順位確率スコアを用いる。重みは、人間の発話交代の時間差分布に合わせた、意図条件付きのタイミング核であり、この指標について有界性、一貫性、二値評価への還元を証明する。11のシステムを比較すると、最良モデルの得点は0.47で、人間の上限値0.86を下回った。このモデルは話者プロファイルにほぼ反応しなかった。TACTと人間の判断とのスピアマン相関は0.81で、二値指標の0.46を上回った。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-23(UTC)
最新改訂
2026-09-23 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Benchmarks for full-duplex spoken dialogue models score turn-taking with binary fixed-window rules that reward immediate response or silence by completeness of the prior turn. We argue that the appropriateness of a response offset, whether delayed silence or anticipatory overlap, is conditional on the speaker's latent intent, identifiable only from that speaker's behavior. We introduce TACT, a benchmark of 9,728 episodes and 73.2 hours from five dyadic corpora; each episode carries dialogue history, a per-speaker memory profile, and an annotator-derived posterior over six intent classes. Scoring replaces binary windows with a strictly proper threshold-weighted continuous ranked probability score whose weights are intent-conditioned timing kernels fitted to human floor-transfer-offset distributions, proving boundedness, consistency, and binary reduction. Across eleven systems the best model reaches 0.47 against a human topline of 0.86, is nearly invariant to speaker profiles, and TACT agrees with human judgments at Spearman 0.81 versus 0.46 for binary metrics.

著者のコメント

Accepted to IEEE SLT 2026

arXiv ID: 2609.27372 / 要約の誤りについて