敵対的学習で音声偽造の検出手順を頑健にする
Robust Workflow Generation via Adversarial Learning for Audio Deepfake Detection
この論文をやさしく読む
ひとことで言うと
偽音声の判定に使う複数の道具を、音声の乱れ方に応じて選び直す仕組みです。
何に役立つ?
雑音や劣化で性能が落ちやすい音声ディープフェイク検出を改善する方法として役立ちます。複数データセットと現実的な音声劣化で比較しています。
この研究の面白いところ
音声を乱すエージェントと、検出道具を選ぶエージェントを対抗させて学習します。単一の検出器を改善するだけでなく、処理の順番や組み合わせそのものを最適化します。
どこまで分かった?
頑健性と汎化の改善は実験で報告されていますが、要旨には具体的な数値や各劣化条件はありません。現実のあらゆる偽音声を確実に検出できることまで示してはいません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
音声合成・声質変換技術の急速な進歩により、音声ディープフェイクはますます現実的になり、実際の応用で深刻なセキュリティ上のリスクを生んでいる。既存の検出手法は、統制された条件では高い性能を示すが、現実世界の摂動や劣化に対しては、しばしば汎化に失敗する。 本論文では、複数の検出ツールを連携させ、頑健な検出ワークフローを動的に構築するフレームワークROGUEを提案する。ROGUEはワークフロー生成を逐次的な意思決定問題として定式化し、2つのエージェントを使う枠組みを導入する。摂動エージェントが音声への摂動を生成し、方策エージェントが摂動のある条件下で検出ツールを選択・実行することを学ぶ。敵対的学習を通じて、摂動を考慮したツール選択、適応的な実行戦略、分布変化に対する頑健性の向上を可能にする。 複数のデータセットと現実世界の劣化を用いた広範な実験で、ROGUEは頑健性と汎化性能の両方において、強力なベースラインを一貫して上回った。この結果は、現実の運用環境で信頼できる音声ディープフェイク検出システムを構築するうえで、敵対的に最適化されたワークフロー生成が有効であることを示している。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-17(UTC)
- 最新改訂
- 2026-09-17 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-17 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
The rapid advancement of speech synthesis and voice conversion technologies has made audio deepfakes increasingly realistic, posing serious security risks in practical applications. While existing detection methods achieve strong performance under controlled conditions, they often fail to generalize under real-world perturbations and corruptions. In this paper, we propose ROGUE, a framework that dynamically constructs robust detection workflows by orchestrating multiple detection tools. ROGUE formulates workflow generation as a sequential decision-making problem and introduces a dual-agent paradigm, where a perturbation agent generates audio perturbations and a policy agent learns to select and execute detection tools under perturbed conditions. Through adversarial learning, ROGUE enables perturbation-aware tool selection, adaptive execution strategies, and improved robustness to distribution shifts. Extensive experiments across multiple datasets and real-world corruptions demonstrate that ROGUE consistently outperforms strong baselines in both robustness and generalization. Our results highlight the effectiveness of adversarially optimized workflow generation for building reliable audio deepfake detection systems in real-world deployment settings.
arXiv ID: 2609.20063 / 要約の誤りについて