視覚情報を学習時だけ使う親指動作のウェアラブル認識
ViMoWear: Visual Motion-Guided sEMG-IMU Representation Learning for Subject-Independent Thumb Gesture Recognition
この論文をやさしく読む
ひとことで言うと
親指のジェスチャーを、学習時は手の三次元動作を使い、利用時はウェアラブルセンサーだけで認識する。
何に役立つ?
義手や拡張現実の操作で、人が変わっても使えるジェスチャー認識に役立つ可能性がある。
この研究の面白いところ
視覚情報を推論時に不要としながら、被験者間の対照学習と細かな動作再構成に利用した。
どこまで分かった?
一人を除く評価で改善したが、要旨には参加人数や改善率の具体値、実運用での評価はない。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
ウェアラブルセンサーを使う手のジェスチャー認識は、人とコンピューターの操作、拡張現実、義手の制御を直感的に行える可能性がある。しかし、ウェアラブル信号は手の動きを間接的に観測するうえ人による差が大きく、学習していない人の動作を認識することは難しい。視覚情報は認識を改善できるが、推論時にも必要とするとセンサー構成が複雑になり、実際の導入を妨げる。本研究は、同期した三次元の手の動きを学習時だけの教師情報として用い、分類時にはウェアラブルセンサーだけを必要とする、視覚的動作に導かれた枠組みViMoWearを提案する。 具体的には、動作に導かれた被験者間の対照学習(MGCL)によって人が変わっても頑健な表現を促し、親指を考慮したマスク付き動作再構成(TMMR)によって細かな動きの情報を保つ。表面筋電図、慣性計測、姿勢が同期したデータ集合で、一人を除いて学習する実験を行うと、複数のセンサー構成にわたり、教師あり学習の比較手法より一貫して改善した。学習した表現は分類器を使わない検索にも対応した。学習時だけ視覚的動作を教師として使うことで、見たことのない人へのウェアラブル表現の一般化を改善した。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-23(UTC)
- 最新改訂
- 2026-09-23 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-23 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Wearable sensing enables intuitive hand gesture recognition for human--computer interaction, augmented reality, and prosthetic control, yet subject--independent recognition remains challenging because wearable signals provide only indirect and highly subject-specific observations of hand motion. Although visual information can improve wearable gesture recognition, requiring it during inference increases sensing complexity and limits practical deployment. We propose ViMoWear, a visual-motion-guided framework that leverages synchronized 3D hand motion as training-only supervision while requiring only wearable sensing for gesture classification at inference. Specifically, Motion-Guided Cross-Subject Contrastive Learning (MGCL) promotes subject-robust representations, and Thumb-Aware Masked Motion Reconstruction (TMMR) preserves fine-grained motion information. The leave-one-subject-out experiments on a synchronized sEMG--IMU--pose dataset demonstrate consistent improvements over supervised baselines across multiple sensing configurations, while the learned representations also support classifier-free retrieval. The proposed training-only visual motion supervision improves the generalization of wearable representations to unseen subjects.
arXiv ID: 2609.27595 / 要約の誤りについて