arXiv論文メモ
新着一覧
cs.CL / cs.HC · 査読状況未確認

長い音声対話の過去情報を使えるか測る評価セット

Vox-Infinity: Benchmarking the Limits of Long-Context Spoken Language Models

Xize Cheng, Wenxu Jia, Chenyuhao Wen, Dongjie Fu, Zehan Wang, Xinyu Zhang, Tao Jin

この論文をやさしく読む

ひとことで言うと

長い音声対話で、以前に話した内容をモデルが覚えて使えるかを測るベンチマークです。会話の往復数と、一回の発話の長さを別々に変えます。

何に役立つ?

音声アシスタントが、質問の直前だけでなく古い情報を根拠に回答できるかを比較するのに役立ちます。必要な根拠の位置が注釈されているため、文脈の長さと性能の関係を調べられます。

この研究の面白いところ

単に長い音声を入力するだけでなく、その質問にどの程度前の情報が必要かを整理しています。7モデルの評価では、根拠が古いほど利用が難しくなる共通傾向が示されています。

どこまで分かった?

要旨には各モデル名、正答率、音声の最大時間の数値がありません。報告はベンチマークでの評価であり、すべての実際の音声対話で同じ失敗が起きると断定するものではありません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

長い文脈の理解は、大規模言語モデルにとって依然として根本的な課題である。入力が長すぎると、モデルは重要な情報を忘れることが多い。この問題は音声領域でさらに顕著になる。音声は圧縮度の低いモダリティであり、意味内容と音響的な手がかりの双方を保つには、テキストより大幅に多くの埋め込みが必要だからである。 この課題に対応するため、音声言語モデルの長文脈理解を評価することに特化して設計された初のベンチマークVox-Infinityを導入する。Vox-Infinityは、対話ターン数とターンの長さという二つの軸で音声履歴を系統的に延長する。対話構造と意味的な複雑さが異なる、多様な代表的場面を網羅する。特に、回答の根拠がどこにあるかを明示する注釈を備え、各質問の解決に必要な過去文脈の量に応じてサンプルを整理することで、長さを考慮した精密な評価を可能にする。 代表的な7つの音声言語モデルを広範に評価した結果、全体として明確な新近性効果が見られた。回答を支える証拠が質問の近くにある場合には、モデルの正確さは一般に高いが、対話履歴のより遠い過去にある証拠を探し出して利用することには苦戦する。事例とデータセットはhttps://vox-infinity.github.ioで公開されている。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-18(UTC)
最新改訂
2026-09-18 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Long-context understanding remains a fundamental challenge for large language models, as excessively long inputs often lead models to forget salient information. This issue is even more pronounced in the speech domain, where audio, as a low-compression modality, requires substantially more embeddings than text to preserve both semantic content and acoustic cues. To address this challenge, we introduce \textbf{Vox-Infinity}, the first benchmark specifically designed to evaluate long-context understanding in spoken language models. Vox-Infinity systematically extends audio history along two dimensions: turn count and turn duration. It covers a diverse range of representative scenarios with varying interaction structures and semantic complexity. Crucially, Vox-Infinity provides explicit answer-provenance annotations and organizes samples according to the amount of historical context required to resolve each query, enabling precise and length-aware evaluation. Extensive evaluations of seven representative spoken language models reveal a clear overall recency effect: models generally achieve higher accuracy when answer-supporting evidence is closer to the query, but struggle to retrieve and use evidence located farther back in the dialogue history. Cases and datasets are available at https://vox-infinity.github.io.

arXiv ID: 2609.22452 / 要約の誤りについて