arXiv論文メモ
新着一覧
cs.CV · 査読状況未確認

顔の向きに応じて唇読みに使う特徴を調整

Pose Adaptive Dynamic FiLM Modulation for Visual Speech Recognition

Matthew Kit Khinn Teng, Haibo Zhang, Takeshi Saitoh

この論文をやさしく読む

ひとことで言うと

顔の向きが変わる唇読みで、特徴の調整を入力に合わせて強くしたり弱くしたりする方法。

何に役立つ?

視覚的音声認識で、複数の姿勢対応経路を組み合わせる際の性能低下を抑える参考になる。

この研究の面白いところ

重みを付けない複数経路より誤り率を下げ、頭の向きの変化が大きいほど深い経路を重視する傾向を示した。

どこまで分かった?

LRS3では提案法の23.91%は単独ResFiLMの20.96%より高い。改善は重みを付けない複数経路との比較である。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

視覚的な音声認識では頭の向きが変わると見た目が大きく変化するため、姿勢に応じた特徴の調整が有用である。ただし、調整の強さが固定された多数の特徴単位線形変調(FiLM)経路を使うと、性能が下がったり特徴が望ましくない形で干渉したりする。著者らは、入力に応じた重みを予測して、姿勢条件付きの調整強度を適応的に制御するDynamic Residual FiLM(DR-FiLM)を備えた枠組みを提案する。LRS2とLRS3での実験では、重みを付けずに複数経路を使うと音素誤り率(PER)がそれぞれ20.33%と29.42%へ悪化した。単独のResFiLMではそれぞれ16.20%と20.96%だった。一方、動的なDeep-Res重みを持つ提案DR-FiLMではPERがLRS2で15.74%、LRS3で23.91%となり、重みを付けない場合の悪影響を大幅に抑えた。学習された重みの分析では、頭の向きの変化が大きいほど、より深いFiLM経路の重みが増す一貫した傾向が見られた。結果は、姿勢条件付きのFiLM経路を組み合わせる際、調整強度を動的に制御する方が効果的だと示す。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-24(UTC)
最新改訂
2026-09-24 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Head-pose variation introduces substantial appearance transformations in visual speech recognition (VSR), making pose-aware feature modulation desirable. However, performance degradation and unwanted feature interactions may result from using numerous Feature-wise Linear Modulation (FiLM) circuits with fixed modulation intensity. We propose a Pose Adaptive Dynamic FiLM framework with a Dynamic Residual FiLM (DR-FiLM) modulator that predicts input-dependent weights to adaptively control the strength of pose-conditioned modulation. Experiments on LRS2 and LRS3 demonstrate that unweighted multi-pathway modulation substantially degrades phoneme recognition, increasing PER to 20.33% and 29.42%, respectively, compared with 16.20% and 20.96% for the single ResFiLM configuration. In contrast, the proposed DR-FiLM with dynamic Deep-Res weighting reduces PER to 15.74% on LRS2 and 23.91% on LRS3, substantially mitigating the adverse effects of unweighted modulation. The analysis of the learned weights further reveals a consistent tendency to assign greater weight to the deeper FiLM pathway as head-pose variation increases. These results show that merging pose-conditioned FiLM circuits is more efficient when the modulation strength is dynamically controlled.

著者のコメント

Submitted for conference publication and currently under review

arXiv ID: 2609.29443 / 要約の誤りについて