arXiv論文メモ
新着一覧
eess.AS · 査読状況未確認

音声の内容や話者らしさを保ちながら訛りの変換を調整する

Partial Accent-Control Editing in Frozen Speech Representations for Accent Conversion

Yangyang Qu and Michele Panariello and Massimiliano Todisco and Nicholas Evans

この論文をやさしく読む

ひとことで言うと

音声の内部表現を編集して訛りを変え、その変更の強さを重みで調整する手法です。専用のアクセント条件付き生成器を新たに学習しません。

何に役立つ?

訛りをどれだけ変えると内容や声の特徴がどれだけ保たれるか、比較・分析する用途が考えられます。要旨では5指標を使った既存手法との比較を報告しています。

この研究の面白いところ

変換を一度に強く適用するのではなく、元音声と参照特徴の融合で強さを調節します。同じ内容を話した対の録音を参照として必要としない点も特徴です。

どこまで分かった?

話者類似度には代償があり、あらゆる声の特徴を完全に保持するという結果ではありません。要旨には評価言語、アクセントの組み合わせ、指標の具体的な数値は示されていません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

アクセント変換とは、言語的な内容や話者に関係するその他の特徴を保持しながら、音声録音が目標のアクセントに近く聞こえるように変更する課題である。多くのアクセント変換システムは学習済みの生成モデルを使う。目標のアクセントを持つ音声を生成できる一方、推論時にアクセント変更の強さを制御できないため、変換の強さが元音声の特徴の保持にどう影響するかを解析しにくい。 本研究では、アクセントを条件とする生成器を学習させず、固定されたWavLM表現の編集に基づいてアクセントを変換する枠組みPartial Accent-Control Editing(PACE)を提案する。まず元音声のWavLM特徴に制約付きの編集を施し、次に、発話内容が対応していないアクセント例から検索した、目標アクセントの参照特徴と融合する。融合の重みを使い、アクセント変換の強さと、元音声のその他の属性が劣化する程度とのトレードオフを制御する。5種類の評価指標を用い、アクセントと絡み合った話者類似度を代償としつつ、PACEがアクセント変換と元音声の保持の両面で、競争力のある2つのベースラインを大幅に上回ることを示す。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-18(UTC)
最新改訂
2026-09-18 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Accent conversion is the task of modifying a speech recording so that it sounds closer to a target accent while preserving linguistic content and other speaker-related characteristics. Most accent conversion systems use trained, generative models. Although they can induce target-accented speech, the strength of accent modification is not controllable at inference time, making it difficult to analyse how the strength of accent conversion affects source preservation. We propose Partial Accent-Control Editing (PACE), an accent conversion framework based upon the editing of frozen WavLM representations without the training of an accent-conditioned generator. A constrained edit is first applied to source WavLM features, which are then fused with target-accent reference features retrieved from non-parallel accent examples. Fusion weights are used to control the trade-off between accent conversion strength and the degradation of other source attributes. Using a suite of five metrics, and with the cost of accent-entangled speaker similarity, we show that PACE is substantially superior to a pair of competitive baselines in terms of both accent conversion and source preservation.

arXiv ID: 2609.22031 / 要約の誤りについて