話者の解剖学的な違いに合わせて声道形状を推定
Anatomy-aware cross-speaker adaptation of complete vocal-tract acoustic-to-articulatory inversion
この論文をやさしく読む
ひとことで言うと
音声から予測した声道の形を、椎骨や歯の目印を使って別の話者の体に合わせる方法です。
何に役立つ?
話者ごとにモデルを再学習せず、声道形状を推定する際の個人差の補正に役立つ可能性があります。
この研究の面白いところ
母音/u/の1画像で目印を定め、その変換を他の録音に再利用し、8話者で位置合わせを評価しました。
どこまで分かった?
評価は別データベースの8話者に基づきます。要旨に臨床的な有用性や、さらに多様な話者での精度は記されていません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
音声から調音器官の形を逆推定するモデルを別の話者に使うには、話者間の解剖学的な違いを考慮する必要がある。著者らは、主に椎骨と歯の構造にある解剖学的な目印を使い、固定した逆推定モデルの予測を未知の話者へ移す幾何学的な適応法を提案する。アフィン変換の後に薄板スプライン変形を施し、予測した声道内の10構造の輪郭を、モデルを再学習せずに各対象話者の形へ写す。 目印は、話者ごとに選んだ母音/u/の1画像で特定し、共通の音韻的基準とする。ただし話者間で調音の形が同じとは仮定しない。得られた写像は各録音を通じて再利用する。単一話者のリアルタイムMRIデータベースでモデルを学習し、別の複数話者MRIデータベースの8話者で適応を評価した。12個または14個の目印を使うアフィン変換と薄板スプラインの構成を比較し、Affine12+TPS14が、点から最も近い点までの平均誤差3.19ミリメートルで最良だった。解剖学的な目印の情報と非剛体の位置合わせを組み合わせる価値を支持する結果である。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-24(UTC)
- 最新改訂
- 2026-09-24 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-24 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Cross-speaker acoustic-to-articulatory inversion requires accounting for anatomical differences between speakers. We propose a geometric adaptation framework that uses anatomical landmarks, primarily on vertebrae and dental structures,to transfer predictions from a fixed inversion model to unseen speakers. An affine transformation followed by thin-plate spline (TPS) deformation maps the predicted contours of 10 vocal-tract structures into each target speaker's geometry without retraining. Landmarks are identified in one selected /u/ frame per speaker as a common phonetic reference without assuming identical articulatory configurations across speakers, and the resulting mapping is reused across recordings. We train the model on a single-speaker rt-MRI database and evaluate adaptation on eight speakers from a separate multi-speaker rt-MRI database. We compare affine and TPS configurations using 12 or 14 landmarks. Affine12+TPS14 achieves the lowest mean point-to-closest-point error of 3.19mm. These results support the combined value of anatomical landmark information and nonrigid alignment.
著者のコメント
Submitted to IEEE ICASSP 2027
arXiv ID: 2609.29766 / 要約の誤りについて