arXiv論文メモ
新着一覧
cs.MM / cs.CV / cs.SD · 査読状況未確認

映像から消した物体の音も取り除くTV-AudioRemover

TV-AudioRemover: Joint Text-Visual Guided Sound Removal with Multi-Task Hard-Mixture Curriculum

Xinyue Guo, Jianxuan Yang, Daiguo Zhou, Jiagao Hu, Yuxuan Chen, Fei Wang, Jian Luan

この論文をやさしく読む

ひとことで言うと

映像から削除した物体に対応する音を、編集済み映像と文章指示を手がかりに混合音から除く方法。

何に役立つ?

映像編集後の音と画の食い違いを減らす用途が考えられる。要旨ではベンチマーク実験での性能を報告している。

この研究の面白いところ

100万規模の映像・音対応データを作り、意味が似た音源を混ぜる難例学習を行い、音の除去を評価するベンチマークも整えた。

どこまで分かった?

最高水準という比較は要旨に記された評価条件での結果。実際のすべての編集映像で自然に聞こえるかは要旨から分からない。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

映像のフレームから対象物を消しても、その音が残れば映像と音が食い違う。既存の映像補完モデルは画素だけを扱い、特に音の除去に使う音声編集モデルは通常テキスト指示に依存するため、意味や時間的な対応付けを助ける複数の情報源を十分に使えない。本研究は、編集済みの映像と自然言語の指示を併用し、元の混合音から消した物体に対応する音を抑えるTV-AudioRemoverを提案する。質の高い学習データを得るため、単一物体の映像と音が対応する100万規模のデータ集合を作る工程を設計し、学習用の混合音と対象音の対を合成する。映像の文脈と指示の意図を利用するため、タスク用トークン、一般化可能な指示モデル、モダリティ別の大域的な誘導をモデルに加える。さらに、複数タスクでの学習によって役割の理解を強め、意味的に似た音の混合を使う難例カリキュラムで微調整し、細かい音源の識別を改善する。評価用に、映像と音からの物体除去ベンチマークAV-Remove-Bench、専用の客観指標、マルチモーダル大規模言語モデルに基づく評価手順も提示する。実験では、主観指標と客観指標の双方で従来の最高水準の性能を達成したとしている。プロジェクトページは https://yjx-research.github.io/TV-AudioRemover/ 。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-22(UTC)
最新改訂
2026-09-22 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Visual object removal can eliminate a target from video frames, yet its acoustic trace persists in the soundtrack, causing obvious audio-visual inconsistency. Existing video inpainting models operate solely on pixels, while audio editing models, especially for the sound removal task, are typically driven by text and therefore rely on limited single-modal control, which is less effective than multimodal guidance that provides stronger semantic grounding and temporal synchronization cues. In this paper, we present Text-Visual Guided Sound Removal (TV-AudioRemover), a target sound removal framework that leverages the visually edited video together with a natural-language instruction to suppress the sound associated with the removed visual object from the original audio mixture. To acquire high-quality training data, we devise a pipeline to construct a million-scale dataset of single-object audio-visual aligned samples, from which we synthesize mixture-target pairs customized for model training. To effectively leverage visual context and follow instruction intent, we augment the model architecture with task tokens, generalizable instruction modeling, and modality-specific global guidance. We further adopt multi-task training to strengthen task-role comprehension, and employ a hard-mixture curriculum that leverages semantically similar acoustic mixtures during fine-tuning to enhance fine-grained source discrimination. To support evaluation, we present AV-Remove-Bench, a comprehensive audio-visual object removal benchmark, along with dedicated objective metrics and an MLLM-based evaluation protocol. Experiments demonstrate that our method achieves state-of-the-art performance on both subjective and objective metrics. Project page: https://yjx-research.github.io/TV-AudioRemover/.

arXiv ID: 2609.25864 / 要約の誤りについて