arXiv論文メモ
新着一覧
cs.SD · 査読状況未確認

顔の動きと音声から狙った歌手の声を小型モデルで分離

MambaVoice: Lightweight Audiovisual Singing Voice Separation Via A Hybrid Mamba-Transformer Model

Adithi Shankar, Gopika Krishnan, Gloria Haro, Xavier Serra, Martín Rocamora

この論文をやさしく読む

ひとことで言うと

映像に映る歌手の顔の動きを手掛かりに、伴奏や別の歌手が混ざった音から、その人の歌声を取り出す方法です。

何に役立つ?

考えられる用途は、音楽映像の歌声抽出を、比較的小さいモデルで処理することです。より大きな映像・音声処理の一部として使う可能性も述べています。

この研究の面白いところ

顔の情報を音声表現のどこに反映するかをゲートで調整し、注意機構と状態空間モデルを組み合わせて時間的な関係を扱っています。

どこまで分かった?

14.18 dBはAcappellaでの音源分離指標SDRで、正答率ではありません。要旨にはURSingの具体値、知覚評価の参加者数、実機での処理遅延は示されていません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

音楽映像から特定の歌声を取り出すことは、複数の歌手や密な楽器伴奏がある場合には特に難しい。特定歌声の分離に向けて、MambaとTransformerを組み合わせた構造を利用する軽量な視聴覚の枠組みMambaVoiceを提案する。モデルは、注意機構に基づく帯域分割音声エンコーダーと、顔の動きの特徴を扱う時空間グラフ畳み込みネットワーク(ST-GCN)によって、音声と映像の系列を共同で符号化する。乗法的ゲーティング機構で両モダリティを融合し、視覚的手掛かりが音声表現を選択的に調節できるようにする。融合した特徴は、Transformerの自己注意と選択的状態空間モデル(SSM)を組み合わせた基盤ネットワークで処理し、線形の計算量で長時間にわたる関係を効率的にモデル化する。 AcappellaとURSingのデータセットで、妨害する歌手の声が混ざる場合を含む難しい条件で評価した。モデルのパラメータ数は1,620万で、AcappellaではSDRが14.18 dBとなり、URSingでもデータセットをまたぐ高い性能を示した。より大きなモデルの一部のパラメータ数で、それらに匹敵する性能を達成している。この結果は、拡張性と効率性を持つ視聴覚音源分離におけるSSMと注意機構の混合構造の有効性を示し、より大きな処理パイプラインの軽量な構成要素に適していることを示唆する。知覚評価も実施し、客観指標の改善をさらに支持する結果を得た。実装をオンラインで公開する。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-22(UTC)
最新改訂
2026-09-22 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Isolating a target singing voice from a music video remains challenging, particularly in the presence of multiple vocalists and dense instrumental accompaniment. We propose MambaVoice, a lightweight audiovisual framework that leverages a hybrid Mamba--Transformer architecture for targeted singing voice separation. The model jointly encodes audio and visual streams using an attention-based band-split audio encoder and a spatio-temporal graph convolutional network (ST-GCN) for facial motion features. These modalities are fused through a multiplicative gating mechanism, enabling visual cues to selectively modulate audio representations. The fused features are processed by a hybrid backbone that combines Transformer self-attention with Selective State Space Models (SSMs), achieving efficient long-range temporal modeling with linear complexity. We evaluated MambaVoice on the Acappella and URSing datasets under challenging conditions, including mixtures with interfering singers. At 16.2 million parameters, the model demonstrates comparable performance, achieving 14.18 dB SDR on Acappella and strong cross-dataset performance on URSing, comparable to larger models at a fraction of the parameter count. These findings highlight the effectiveness of hybrid SSM--attention architectures for scalable, efficient audiovisual source separation, suggesting they are well-suited as lightweight components within larger pipelines. We conduct a perceptual study that further supports our improvements in objective metrics. We provide our implementation online.

著者のコメント

6 pages, Accepted at ISMIR 2026

arXiv ID: 2609.26635 / 要約の誤りについて