arXiv論文メモ
新着一覧
cs.CL / eess.AS · 査読状況未確認

会話の流れを取り入れる音声感情認識

Enriching Speech Emotion Representations with Conversational Context

Arthur Peuvot, Romaric Besançon, Gaël de Chalendar, Bianca Vieru and Ioana Vasilescu

この論文をやさしく読む

ひとことで言うと

一つの発話だけでなく、その前後の会話の流れを取り込んで音声から感情を認識する方法です。

何に役立つ?

会話型音声インターフェースで、発話の背景にある感情の推移を利用する設計に役立つ可能性があります。要旨で示されたのはデータセット上の認識性能です。

この研究の面白いところ

長さを変えられる文脈窓を用い、性能向上の原因が話者の識別や録音条件ではなく、感情と会話の連続性にあることを要素除去実験で調べています。

どこまで分かった?

報告された性能はIEMOCAP、SAFE、MELDと、クラスごとに均等に扱う指標での結果です。実際の対話システムの利用者への効果は要旨には示されていません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

感情の検出は、人と正確かつ柔軟にやり取りできるシステムを構築するうえで必要である。音声感情認識は知的な音声インターフェースの開発で重要な研究対象となっている。しかし多くの研究は発話単位で感情を予測し、感情の推移や話者間のやり取りを含む会話の文脈を無視している。本論文では、さまざまな長さの会話文脈を取り込み、音声でのやり取りの中で感情がどう変わるかをよりよく捉えるモジュールACERT(時間を通じて平均した文脈的感情表現)を導入する。この方法の頑健性を評価するため、感情表現の様式と文脈が多様なデータセットで実験した。クラスごとに均等に評価する重みなしの指標で、ACERTはIEMOCAPにおいて現行の最先端手法を上回り、SAFEでは文脈を考慮した初めての基準結果を設定し、MELDでは良好な結果を得た。要素を取り除く実験から、改善は話者の識別情報や音響条件ではなく、感情と会話の連続性によることが示された。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-22(UTC)
最新改訂
2026-09-22 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Detecting emotions is necessary for building systems that can accurately and adaptively interact with humans. Speech Emotion Recognition (SER) has become an important research focus to develop intelligent spoken interfaces. However, most studies predict emotions at the utterance level, ignoring the conversational context, along with the emotional flow and speaker interactions it carries. In this paper, we introduce ACERT (Averaged Contextual Emotion Representation through Time), a module that integrates a flexible-length window of conversational context to better capture emotional evolution in spoken interactions. To evaluate the robustness of this method, we conducted experiments on datasets spanning diverse emotionally expressive styles and contexts. ACERT outperforms current state-of-the-art (SOTA) approaches on IEMOCAP, establishes the first context-aware benchmark on SAFE, and obtains strong results on MELD for unweighted, class-balanced metrics. Ablation studies show that ACERT's gains come from emotional and conversational continuity, rather than from speaker identity or acoustic conditions.

著者のコメント

5 pages, 1 figure, 2 tables. Submitted to ICASSP 2027

arXiv ID: 2609.26422 / 要約の誤りについて