都市の音と映像をまとめて説明するデータセット
AVSD-Scenes: A Dataset for Audio-Visual Description of Urban Scenes
この論文をやさしく読む
ひとことで言うと
都市の音声と映像をそれぞれAIで説明し、その文章を統合した12,291件のデータセットです。音だけ、映像だけでは得られない情報を文章にまとめます。
何に役立つ?
都市の音や映像を文章で検索したり、情景を分類したりする手法の評価に役立ちます。説明による分類で最大94.5%、複数の埋め込みを合わせた分類で95.4%が報告されています。
この研究の面白いところ
説明生成に情景ラベルを与えない場合も調べています。ラベルを書き写しただけで分類が容易になったのか、それとも音と映像から内容を捉えたのかを確かめる評価です。
どこまで分かった?
既存のTAUデータセットを元にAIが生成した説明を評価しています。要旨には外部の都市データへの汎化や、個々の説明の誤記率は示されていません。分類正解率は説明の全記述が正しい割合とは異なります。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
自然言語による説明は、音と映像からなる都市の情景を意味の豊かな表現に変えられるが、聴覚情報と視覚情報を一緒に記述するデータセットはまだ限られている。本論文では、都市環境を対象とした、音と映像のペアに対応する情景説明データセットAVSD-Scenesを導入する。このデータセットには、TAU Urban Audio-Visual Scenesデータセットから生成した12,291件の音と映像の情景説明が含まれる。 構築にあたり、まずQwen2-Audio-7BとQwen2.5-VL-7Bをそれぞれ使い、音に基づく説明と映像に基づく説明を生成する。次に、Qwen3-14B、Mistral-Small-3.2-24B-Instruct-2506、Gemma-3-27B-itという大規模言語モデルを用いて、各モダリティに固有の説明を組み合わせ、双方の相補的な情報を捉えたマルチモーダルな説明を生成する。AVSD-Scenesを、意味の整合性、異なるモダリティ間の検索、情景分類、LLMを判定役とする評価、人による主観評価によって検証する。 結果は、マルチモーダルな説明が、モダリティごとに作った説明と比べ、情景を識別する情報を十分に保ちながら、意味の整合性と異なるモダリティ間の検索性能を改善することを示す。生成した説明は都市の情景分類で最大94.5%の正解率に達し、音・映像・説明の埋め込みを組み合わせると、正解率はさらに95.4%へ向上する。また、プロンプトの指示から情景ラベルを取り除いても、説明は高い情景識別性を維持する。これは、説明が単にラベル情報を反映しているのではなく、音と映像の内容に由来する意味情報を捉えていることを示している。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-10-01(UTC)
- 最新改訂
- 2026-10-01 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-10-01 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Natural language descriptions can provide rich semantic representations of audio-visual urban scenes, yet datasets that jointly describe both auditory and visual information remain limited. In this paper, we introduce AVSD-Scenes, a paired audio-visual scene description dataset for urban environments. The dataset contains 12,291 audio-visual scene descriptions generated from the TAU Urban Audio-Visual Scenes dataset. To construct the dataset, we first generate audio- and visual-based descriptions using Qwen2-Audio-7B and Qwen2.5-VL-7B, respectively. These modality-specific descriptions are then combined using large language models, namely Qwen3-14B, Mistral-Small-3.2-24B-Instruct-2506, and Gemma-3-27B-it, to produce multimodal descriptions that capture complementary information from both modalities. We benchmark AVSD-Scenes using semantic alignment, cross-modal retrieval, scene classification, LLM-as-a-judge evaluation, and human subjective assessment. Results show that multimodal descriptions improve semantic alignment and cross-modal retrieval performance compared with modality-specific descriptions while preserving strong scene-discriminative information. The generated descriptions achieve up to 94.5% accuracy in urban scene classification, while combining audio, visual, and description embeddings further improves accuracy to 95.4%. Furthermore, the descriptions remain highly scene-discriminative even when scene labels are removed from the prompting instructions, indicating that they capture semantic information derived from the audio-visual content rather than merely reflecting label information.
著者のコメント
Submitted to ICASSP 2027
arXiv ID: 2610.01861 / 要約の誤りについて