arXiv論文メモ
新着一覧
cs.CL · 査読状況未確認

少ないデータで音声と文章の表現をグラフで揃える

Low resource cross-modal alignment using HGNN to enhance speech representation

Yannick Yomie Nzeuhang, Marie Tahon, Paulin Melatagia Yonta

この論文をやさしく読む

ひとことで言うと

音声と文章の対応を、少ないデータと計算資源で学ぶグラフ手法を提案する。

何に役立つ?

低資源言語の単語検索や音声表現学習の方法を選ぶ参考になる。

この研究の面白いところ

Yembaの単語検索では、少ない資源で比較手法SAMU-XLSRを上回った。

どこまで分かった?

評価は英語TIMITとYembaの単語単位の課題に基づく。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

音声と文章の表現空間を揃えることは、異なる音声と文章を共通の表現空間へ写し、それぞれの情報を豊かにする複数種類の情報を扱う学習法である。SAMU-XLSRなどの既存の構成は、通常、生徒・教師の枠組みで音声エンコーダーを追加学習し、文章の表現に近い音声表現を作る。これにより音声表現に意味情報が加わる。しかし、この種のシステムは一般に大量の学習データと計算資源が必要で、資源の限られた低資源言語への適用は難しい。 本研究は、異種グラフニューラルネットワークとリンク予測に基づく、データ効率のよい空間整合法を提案する。中心となるのは、メッセージ伝播によって文章から音声へ情報を明示的に移し、大規模な学習データの必要性を減らしながら、音響表現を本質的かつ解釈しやすい形で豊かにするという考えである。高資源言語では十分に研究されているものの、単語単位の音声課題は一部の低資源言語でなお重要である。そこで、英語のTIMITとカメルーンの言語Yembaを使い、単語単位の音声・文章の対応付けを実験した。提案法は先進的なSAMU-XLSRと同程度の結果を出し、Yembaでの単語検索ではそれを上回った。使用資源は大幅に少なく、方法の能力と省資源性を示す。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-19(UTC)
最新改訂
2026-09-19 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Speech-text space alignment is a multimodal representation learning method consisting to map different speech and text into a shared representation space, leading to enrichment of the representation of each modality. Proposed architectures, such as SAMU-XLSR, typically follow a student/teacher framework, with the goal of fine-tuning an audio encoder to produce representations that closely match those of the text. In this way a speech representation is semantically enriched. However, such systems generally require large amounts of training data and considerable computational resource, making them difficult to apply to low resources languages under frugal constraints. The present work proposes a data-efficient space alignment method based on Heterogeneous Graph Neural Networks and link prediction. The core idea is to leverage message passing to explicitly transfer information from the text modality to the speech modality, thereby reducing the need for large training datasets and intrinsically enriching the acoustic representations, all in a more interpretable manner. Although thoroughly explored for high-resource languages, word-level tasks in speech remain relevant for certain low-resource languages. Therefore, we conducted experiments on speech-text alignment at the word level using the TIMIT (English) dataset and Yemba (a Cameroonian language). Our approach yields results comparable to those of SAMU-XLSR, a state-of-the-art method, and even surpasses it for the Yemba language in the task of word retrieval, while using far fewer resources, demonstrating its power, frugality, and efficiency.

arXiv ID: 2609.23191 / 要約の誤りについて