arXiv論文メモ
新着一覧
cs.CV / eess.IV · 査読状況未確認

会話中の高速内視鏡映像から喉頭の各部位を分割

Laryngeal Structure Segmentation in High-Speed Videoendoscopy Using Deep Learning

Sardar Nafis Bin Ali, Mohsen Zayernouri, Dimitar D. Deliyski, Maryam Naghibolhosseini

この論文をやさしく読む

ひとことで言うと

話している最中の高速内視鏡映像から、声帯や喉頭蓋などを自動で切り分ける研究です。静かに母音を伸ばす場面だけでなく、大きく動く連続発話も対象にしています。

何に役立つ?

考えられる用途は、大量の内視鏡フレームを解析し、発声時の各組織の動きを数値化することです。異常検出への臨床利用は将来の可能性として述べられています。

この研究の面白いところ

正常音声と障害音声、持続母音と連続発話を含むデータで複数の部位を学習しています。画像前処理に加えて、数値指標と目視の両方で評価しています。

どこまで分かった?

95%超は画像分割ネットワークの全体正解率であり、音声障害の診断正答率ではありません。要旨には患者数、部位ごとの分割指標、臨床診断への効果は示されていません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

喉頭の高速ビデオ内視鏡(HSV)は、さまざまな発声条件で、声帯の振動とともに喉頭の各構造の動きを観察する有効な方法である。喉頭組織の領域分割によって、異なる組織構造とその動態を解析でき、発声における喉頭筋の関与を特徴付ける助けとなる。HSVのフレーム数は膨大なため、この作業の自動化は不可欠である。過去の研究でも深層学習による喉頭構造の分割は行われてきたが、連続した発話中のHSVデータには適用されていなかった。このデータには、大きな組織運動と光ファイバー画像取得に伴う画質制限という課題がある。連続発話データへの深層学習の適用は、定常的でない喉頭の振る舞いを捉え、音声障害に関連する異常パターンを見つけるために重要である。 本研究はこの不足に対処するため、正常音声と障害のある音声から得た、母音の持続発声および連続発話のHSVデータを用い、披裂喉頭蓋ひだと披裂軟骨、声帯、喉頭蓋、声門領域を検出するU-Netモデルを学習する。ノイズ除去やヒストグラム平坦化などの画像前処理を施し、学習用HSV画像の品質とネットワークの性能を高めた。最後に、定量的な性能指標とテスト画像の目視による定性的点検を併用し、ネットワークの精度と信頼性を評価した。開発したネットワークは全体の正解率が95%を超える高性能を示し、喉頭画像の自動解析、喉頭動態の定量的特徴付け、将来の臨床現場における異常な喉頭挙動の検出に向け、信頼できる道具となる可能性を示した。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-22(UTC)
最新改訂
2026-09-22 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Laryngeal high-speed videoendoscopy (HSV) offers an effective means of observing the motion of different laryngeal structures along with vibratory behaviors of the vocal folds under various voicing conditions. Segmentation of laryngeal tissues enables analysis of different tissue structures and their dynamics, helping characterize the involvement of laryngeal muscles in voice production. Given the large number of HSV frames, automating this task is imperative. While deep learning-based methods have been implemented in previous studies to segment laryngeal structures, they have not been applied to HSV data during connected speech, which poses significant challenges due to excessive tissue movements and image quality limitations associated with fiberoptic image acquisition. The application of deep learning to connected speech data is critical for capturing nonstationary laryngeal behaviors and identifying anomalous patterns associated with voice disorders. The present study aims to address these gaps by training U-Net models to detect the aryepiglottic folds and arytenoid cartilages, vocal folds, epiglottis, and glottal area, using HSV data from both sustained vowel phonation and connected speech obtained from normophonic and disordered voices. Image pre-processing techniques, including noise removal and histogram equalization, were applied to improve the quality of the training HSV images and enhance network performance. Finally, to evaluate the accuracy and reliability of the networks, quantitative performance metrics were used alongside qualitative visual inspection of the test images. The high performance of the developed networks, with overall accuracies exceeding 95%, establishes their potential as reliable tools for automated laryngeal image analysis, quantitative characterization of laryngeal dynamics, and future detection of anomalous laryngeal behaviors in clinical settings.

著者のコメント

22 pages, 9 figures

arXiv ID: 2609.26636 / 要約の誤りについて