音と映像の教師を役割別に使う出来事の時間位置推定
If You Hear It, Help Find It: Orthogonal Knowledge Distillation for Open-Vocabulary Audio-Visual Event Localization
この論文をやさしく読む
ひとことで言うと
音と映像から出来事の時刻を探す際、境界に強い視覚教師と意味に役立つ音声教師を別の役割で使う。
何に役立つ?
文章で指定した出来事を動画中で見つけるモデルの蒸留設計に役立つ。
この研究の面白いところ
教師の優劣を一般化せず、この評価設定での信頼性に応じて判断経路と補助経路を分けた。
どこまで分かった?
視覚教師が境界に強いという判断はOV-AVEBenchの設定に限る。結果も指定ベンチマークでの評価である。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
開放語彙の音・映像イベント位置推定は、文章で指定された出来事が動画と音声のどの時刻に起きたかを求める。使える教師情報は時間的な境界の信頼性が異なる。OV-AVEBenchで設定した視覚の教師は、設定した音声の教師より境界の手掛かりとして信頼できた一方、音声の教師にも意味的な情報があった。これはこの設定に限る診断であり、視覚と音声の一般的な優劣ではない。そこで、どの教師の情報を位置推定の判断に使い、どれを補助にとどめるかという問題として捉え、信頼性に応じた非対称な蒸留法OV-OrthKDを提案する。視覚の特徴転写は判断に沿った表現を形作り、音声の特徴転写は別の補助的な部分空間を豊かにする。文章の原型表現は既知・未知の種類の意味を固定し、直交損失が二教師の射影方向の重なりを抑える。推論時には質問に応じて両方の情報を融合する一方、標準の訓練では音声教師の監督を時間区間の判定経路へ直接入れない。OV-AVEBenchで区間AP 0.816を達成し、公式の微調整基準よりF1@0.5が全体で2.7ポイント、未知カテゴリで3.4ポイント高かった。経路の割当、教師の役割交換、入力の破損、転移の分析も、教師情報の置き場所がこの課題に固有の設計要素であることを支持した。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-20(UTC)
- 最新改訂
- 2026-09-20 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-20 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Open-vocabulary audio-visual event localization (OV-AVEL) grounds a text-queried event in time from video, audio, and language. The supervision sources available to this task can differ in temporal-boundary reliability: on OV-AVEBench, our configured visual teacher gives more reliable boundary cues than the configured audio teacher, although the latter is a strong pretrained audio model and remains semantically informative. This is a setting-specific diagnostic rather than a universal ranking of vision and audio. We formulate the resulting challenge as supervision placement: which teacher signals may shape the localization decision, and which should remain auxiliary. Based on this view, we propose OV-OrthKD, a reliability-aware asymmetric distillation framework. Visual feature transfer shapes a decision-aligned representation, audio feature transfer enriches a complementary auxiliary subspace, a text prototype anchors seen/unseen category semantics, and an orthogonality loss limits directional overlap between the two teacher-specific projections. The student continues to use both modalities through query-aware fusion at inference, while the default training recipe keeps audio-teacher supervision off the segment-logit path. On OV-AVEBench, OV-OrthKD achieves 0.816 segment AP and improves F1@0.5 over the official fine-tuning baseline by 2.7 points overall and 3.4 points on unseen categories. Path-assignment, role-swap, corruption, and transfer analyses consistently support supervision placement as a task-specific design axis for OV-AVEL.
著者のコメント
Accepted to ACM Multimedia 2026 (poster). 9 pages, 5 figures
arXiv ID: 2609.23376 / 要約の誤りについて