arXiv論文メモ
新着一覧
eess.AS · 査読状況未確認

発話内容が違う参照音声から局所的な声の変換を求める

Shared-State Local Translations for Training-Free Voice Conversion

Yangyang Qu and Michele Panariello and Massimiliano Todisco and Nicholas Evans

この論文をやさしく読む

ひとことで言うと

参照音声と元音声で話している内容が違っても、音声表現を共通の局所状態に分け、状態ごとの差を使って声質を変える方法です。

何に役立つ?

同じ文章を話した対応音声がない状況で、短い参照発話から声質変換を行う用途が考えられます。

この研究の面白いところ

フレーム同士を無理に対応させず、両方の音声をまとめて作った状態の平均差を使います。各フレームへの更新は、その状態に属する確率に応じて変わります。

どこまで分かった?

学習不要という設定でも、音声の組ごとにガウス混合モデルを当てはめます。評価はLibriSpeechの指定手順によるもので、自然さが最高という比較範囲は学習不要システムに限られます。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

1回の参照で行う学習不要の声質変換(VC)では、元音声と参照音声の言語的内容が異なる場合があるため、信頼できるフレーム単位の対応を仮定できない。本研究ではStateVCを提案する。元音声と参照音声のフレーム単位のWavLM表現をまとめ、そこから共通する局所領域の集合を同時に定義する。これらの領域を「状態」と呼ぶ。 共有状態は、元音声と参照音声の明示的なフレーム対応付けを行わず、まとめた表現にその組専用のガウス混合モデルを当てはめることで得る。各状態の中で、StateVCは元のWavLM空間における、元音声から参照音声への平均のずれを推定する。次に、元音声のフレームの事後確率で状態ごとのずれを組み合わせるため、フレームごとに異なる局所的な更新を与えられる。 LibriSpeechの1回参照の評価手順では、StateVCは評価したシステムの中で最も低い単語誤り率(WER)8.01%と文字誤り率(CER)3.22%を達成し、話者類似度(SIM)は0.9512である。また、評価したシステムの中で知覚された話者類似度の平均が最も高く、評価した学習不要システムの中で自然さの平均が最も高い。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-10-01(UTC)
最新改訂
2026-10-01 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

In one-shot training-free voice conversion (VC), the source and reference utterances may contain different linguistic content, so reliable frame-level correspondence between them cannot be assumed. We propose StateVC, which jointly defines a common set of local regions from pooled frame-level WavLM representations of the source and reference utterances; we refer to these regions as states. These shared states are obtained by fitting a pair-specific Gaussian mixture model to the pooled representations, without explicit source--reference frame matching. Within each state, StateVC estimates a source-to-reference mean shift in the original WavLM space. Source-frame posterior probabilities then combine the state-specific shifts so that different frames can receive different local updates. For the LibriSpeech one-shot protocol, StateVC achieves the lowest word error rate (WER) and character error rate (CER) among the evaluated systems, at 8.01% and 3.22%, respectively, with a speaker similarity (SIM) of 0.9512. It also achieves the highest mean perceived speaker similarity among the evaluated systems and the highest mean naturalness among the evaluated training-free systems.

arXiv ID: 2610.01952 / 要約の誤りについて