arXiv論文メモ
新着一覧
eess.AS · 査読状況未確認

少量の音声から音響特徴を精密に変える合成モデルARIS

ARIS: Low-Resource Glass-Box Neural Source-Filter Synthesis for Phonetic Stimulus Manipulation

Yiran Ding, Wenwei Xu

この論文をやさしく読む

ひとことで言うと

少ない音声データで、声の高さや共鳴などを個別に調整した実験用音声を作るモデルです。

何に役立つ?

音声学で特定の音響的手がかりだけを変えた刺激を作る用途が考えられます。要旨では三言語、五つの単一話者データで比較しています。

この研究の面白いところ

ニューラルモデルで係数を推定し、実際の合成は決定論的に行うため、各音響パラメータを直接編集できます。

どこまで分かった?

HiFi-Glotより予測自然さはやや低いと報告しています。評価は小規模な単一話者コーパスでの結果です。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

音声学の研究者は、音声の品質をある程度保ちながら、特定の音響的な手がかりを正確に変えた刺激を作る必要がある。古典的な合成法と現代のニューラル法には、パラメータの精密な制御と高い忠実度の間に兼ね合いがあり、ニューラル合成には通常、研究者が容易には得られない量のデータが必要である。本研究は、ニューラルなパラメータ推定と決定論的なデジタル信号処理による合成を組み合わせた、音源・フィルタ型のニューラルモデルARISを提示する。制御項目はすべて合成器の係数なので、基本周波数F0、フォルマント、声門音源を直接編集できる。 三言語にまたがる五つの小規模な単一話者コーパスで、ARISはWORLDと同程度の品質で音声を再合成し、Praat KlattGridより単一パラメータを正確に編集でき、手がかり間の意図しない影響は無視できる程度だった。大規模コーパスで事前学習し、同じデータで追加学習したHiFi-Glotと比べると、予測された自然さはやや低かったが、元の録音をより忠実に再現し、より精密に変換できた。音声サンプルも著者らのサイトで公開している。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-24(UTC)
最新改訂
2026-09-24 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Phoneticians often need to construct stimuli in which specific acoustic cues are precisely manipulated while preserving decent speech quality. Classical synthesis and modern neural methods sit along a trade-off between precise parametric control and high fidelity, and neural synthesis typically demands more data than phoneticians can easily obtain. We present ARIS (Analytic Resonant Interpretable Synthesis), a neural source-filter model that pairs neural parameter estimation with deterministic DSP synthesis. Every control is a coefficient of the synthesizer, so F0, formants and the glottal source can be edited directly. On five small single-speaker corpora in three languages, ARIS resynthesizes speech with quality comparable to WORLD and edits single parameters more accurately than Praat KlattGrid, with negligible crosstalk between cues. Compared with HiFi-Glot, pre-trained on a large corpus and fine-tuned on the same data, ARIS scores slightly lower on predicted naturalness but reproduces the recordings more faithfully and manipulates them more precisely. Audio samples: https://n1r.github.io/ARIS_nsf/.

著者のコメント

5 pages, 2 figures, 3 tables. Submitted to ICASSP 2027. Audio samples: https://n1r.github.io/ARIS_nsf/

arXiv ID: 2609.29923 / 要約の誤りについて