音声と文章の不一致も扱う感情認識モデルBiCFlow-MER
BiCFlow-MER: Orchestrating Discriminative and Generative Multimodal Emotion Recognition via Conditional Transport
この論文をやさしく読む
ひとことで言うと
音声と文章が食い違う場合も、感情の手掛かりを分けて検証するモデルを提案した。
何に役立つ?
音声と文章を使う感情認識システムの設計に役立つ可能性がある。
この研究の面白いところ
感情の証拠を明示的な空間へ運び、往路と逆路の整合性で候補感情を確認する。
どこまで分かった?
三つのベンチマークで比較手法を上回ったが、要旨には差の数値や実運用での評価は示されていない。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
複数の情報源を使う感情認識では、異なる種類の手掛かりを統合して人の感情状態を推定する。音声と文章を使う場合、感情の手掛かりは話者の話し方や語句の内容と絡み合い、情報源間の不一致も判断を難しくする。従来の識別型の融合では、複数の手掛かりが最終予測へ圧縮され、情報源ごとの手掛かりや衝突の情報が十分に残らない。一方、大規模な生成型感情モデルでは、感情に関する推論が言語生成に組み込まれるため、証拠が暗黙のままで、構造化された空間での検証が難しい。 これに対してBiCFlow-MERを提案し、音声・文章の感情認識を、構造化された感情空間内の生成的な証拠の輸送として定式化する。感情に関係する証拠を、話者の話し方や語句の内容から分離して、情報源間の衝突を考慮した感情条件を作る。その条件に導かれ、各発話を双方向の整流フローによって感情空間の明示的な到達点へ運ぶ。候補感情は、到達点に対する適応的なプロトタイプ群のスコアと、元の複数情報源の条件に対する逆方向のクラス整合性を組み合わせて検証する。これにより不一致を考慮した認識が可能になる。IEMOCAP、MELD、ゼロショットのCASEベンチマークの全てで、比較した手法を上回った。条件付き輸送によって識別型の認識と生成型の証拠モデル化を組み合わせる、新たな感情認識の枠組みを示す。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-23(UTC)
- 最新改訂
- 2026-09-23 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-23 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
In multimodal emotion recognition (MER), human affective states are inferred by integrating complementary cues from multiple modalities. In audio-text MER, affective cues are often entangled with speaker style and lexical content, while cross-modal disagreement further complicates how the evidence should be integrated. Under conventional discriminative fusion, multimodal evidence is compressed into a terminal prediction, with modality-specific cues and conflict information insufficiently preserved. In large generative affective models, by contrast, affective reasoning is typically embedded in language decoding, leaving emotion evidence implicit and difficult to verify in a structured space. To address these limitations, BiCFlow-MER (Bidirectional Conditional Flow for Multimodal Emotion Recognition) is proposed as a conditional-flow framework in which audio-text MER is formulated as generative evidence transport within a structured emotion space. Within BiCFlow-MER, emotion-oriented evidence is disentangled from speaker-style and lexical-content factors to construct a conflict-aware affective condition. Guided by this condition, each utterance is transported to an explicit emotion-space endpoint through a bidirectional rectified flow. Candidate emotions are jointly verified through adaptive prototype-cloud scoring of the transported endpoint and backward class-to-condition consistency with the original multimodal condition, enabling conflict-aware recognition. BiCFlow-MER is shown to outperform all compared methods across IEMOCAP, MELD, and the zero-shot CASE benchmark. By orchestrating discriminative recognition and generative evidence modeling through conditional transport, BiCFlow-MER defines a new MER paradigm.
arXiv ID: 2609.27615 / 要約の誤りについて