音声・映像・会話からAIが本当の依頼を理解できるか
Omni Demand Understanding: A Benchmark for Contextual User-Intent Inference in Multimodal Interaction
この論文をやさしく読む
ひとことで言うと
AIが返事の前に、そもそも自分への依頼なのか、何を求められているのかを分かるかを測る評価です。声だけでなく映像や会話の流れが必要な状況を扱います。
何に役立つ?
音声・映像アシスタントの意図の取り違えや、依頼がないのに反応する問題を評価するのに役立ちます。返答の出来と依頼理解の正しさを分けて測れます。
この研究の面白いところ
重要情報の回収率と、要求がない場面の誤作動を両方測っています。評価した14モデル中11モデルで誤作動率が50%を超えたと報告されています。
どこまで分かった?
数値は本研究のODU-Benchと評価対象モデルにおける結果です。44.7%は文脈から推定すべき重要情報の回収率で、すべての一般的な質問への正答率ではありません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
自然な音声・映像によるやり取りは、慎重に書いたテキストプロンプトの代わりに発話と視覚で伝えられる、AIアシスタントの重要なインターフェースになりつつある。しかし、既存の対話能力ベンチマークは依然として主に応答品質に焦点を当てており、より根本的な問い、すなわち複雑なマルチモーダルのやり取りからモデルがユーザーの根底にある要求を正しく推定できるかは十分に探究されていない。現実のユーザー要求は発話だけでは十分に指定されないことが多く、複数のモダリティの手掛かりや対話履歴から推定する必要がある。曖昧または流暢でない表現や騒がしい音響環境が、この推定をさらに難しくする。逆に、依頼らしい発話がアシスタントへの要求ではない場合もあり、誤作動を招く。本研究はOmni Demand Understanding(ODU)を、独立したマルチモーダル文脈推論問題として設定する。やり取りのストリームが与えられたとき、モデルはユーザーの要求があるかを検出し、複数のモダリティと会話の文脈から意図を推定しなければならない。ODUは単一ターンと複数ターンのやり取りを含む五つの側面でこの能力を評価する。課題に基づく分類体系、その分類に沿ったエージェントによる動画生成、人間が記録したやり取りを使い、メディアに根拠を置く注釈付けと人手の検証を経てODU-Benchを構築する。ネイティブなマルチモーダル大規模言語モデル14種を評価した。最も強いGemini 3.1 Proでも、視覚、音響、会話の文脈から推定すべき重要情報の44.7%しか取り出せなかった。さらに14モデル中11モデルで、要求がない場面での誤作動率が50%を超えた。これらの結果は、現在のモデルが文脈からユーザー要求を推定する能力に、系統的な不足があることを示す。ODUが、適切な応答を生成する前にユーザーの要求を正しく理解するという、マルチモーダル対話で不可欠でありながら十分に研究されてこなかった能力の評価を確立することを期待する。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-18(UTC)
- 最新改訂
- 2026-09-18 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-18 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Natural audio-visual interaction is emerging as an important interface for AI assistants, allowing users to communicate through speech and vision rather than carefully composed text prompts. However, existing benchmarks of interactive capabilities still focus primarily on response quality, leaving a more fundamental question underexplored: can a model correctly infer the user's underlying demand from complex multimodal interaction? Real-world user demands are often underspecified in speech and must be inferred from multimodal cues and dialogue history. This inference is further complicated by ambiguous or disfluent expression and noisy acoustic environments. Conversely, request-like speech may not constitute a demand to the assistant, leading to false triggers. We establish Omni Demand Understanding (ODU) as a distinct multimodal contextual inference problem: given an interaction stream, a model must detect whether a user demand is present and infer intent from multimodal and conversational context. ODU evaluates this capability along five dimensions, covering both single-turn and multi-turn interactions. We construct ODU-Bench using a challenge-driven taxonomy, taxonomy-guided agentic video generation, and human-recorded interactions, followed by media-grounded annotation and human verification. We evaluate 14 native MLLMs. Even the strongest, Gemini 3.1 Pro, recovers only 44.7% of key information that must be inferred from visual, acoustic, or conversational context. Moreover, 11 of the 14 models exhibit false-trigger rates above 50% on non-demand scenarios. These results reveal a systematic capability gap in current MLLMs' ability to infer contextual user demands. We hope ODU can establish the evaluation of a previously underexplored yet essential capability in multimodal interaction: correctly understanding user demands before generating an appropriate response.
arXiv ID: 2609.21392 / 要約の誤りについて