音の名前が分からないときにも伝わる認識結果
Misrecognition or Abstraction? Rethinking Outputs of Sound Event Recognition
この論文をやさしく読む
ひとことで言うと
音の正体を言い切れないときも、候補と確信度に擬音語を添えて聞こえ方を伝える方法。
何に役立つ?
環境音の認識結果を人が理解し、周囲の状況を判断するための表示に役立つ。
この研究の面白いところ
認識精度を保ちながら、曖昧な音の特徴を擬音語で補足し、人の聴取評価でも好まれた。
どこまで分かった?
評価は二つの指定データセットと記載の評価方法で行われ、すべての環境音で有効とは示していない。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
一般的な音の認識システムは通常、特定の音の種類を一つのラベルとして出し、入力音から正しい種類を特定できると暗黙に仮定する。しかし実際の音は、種類を明確に特定できないことも多い。人間は種類が分からない曖昧な音からも、周囲の状況を理解できる。そこで、不確実な場合に音認識の出力をどう作り直すべきか検討するため、音の種類、予測の信頼度、その音を表す擬音語を組み合わせた出力表現を提案する。従来の種類別の認識を保ちながら、予測が不確かな場合にも役立つ音響的な特徴を擬音語で伝える。ESC-50とESC-50-Onomatopoeiaを使った実験では、従来の認識専用システムと同程度の音認識性能だった。また、大規模言語モデルを評価者とする試験と人間の聴取実験では、特に周囲の状況を理解する補助として使う場合、音のラベルだけによる従来の確定的な出力より提案形式が好まれた。不確実な音についても情報を伝えやすい表現となる可能性を示す。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-20(UTC)
- 最新改訂
- 2026-09-20 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-20 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Conventional general sound recognition systems typically output deterministic sound event labels, implicitly assuming that the target sound class can be correctly identified from the input audio. However, in real listening situations, the sound event class is not always clearly identifiable. Human listeners may nevertheless understand their surroundings from an ambiguous sound without identifying its exact sound event class. This motivates a discussion of how the outputs of sound recognition systems should be redesigned under such uncertainty. As a basis for this discussion, this paper proposes an output representation for sound event recognition that combines a sound event class, its confidence score, and an onomatopoeic description of the sound. The proposed representation preserves conventional class-based recognition while providing an additional onomatopoeic description of acoustic characteristics that can remain informative even when the class prediction is uncertain. Experiments using ESC-50 and ESC-50-Onomatopoeia show that the proposed method achieves sound recognition performance comparable to that of a conventional recognition-only system. In addition, an LLM-as-a-judge evaluation and subjective listening experiments indicate that the proposed output is preferred over conventional deterministic outputs based on the sound event label, particularly when used to support understanding of the surrounding environment. These results suggest that such output representations can make sound event recognition more informative and communicative under uncertainty.
著者のコメント
Accepted to DCASE 2026 Workshop
arXiv ID: 2609.23411 / 要約の誤りについて