arXiv論文メモ
新着一覧
cs.HC / cs.SD · 査読状況未確認

音を出す動作だけ短く録音する手首センサーの活動認識

AnomaSense: Anomaly-based Sensor Activation for Fine-Grained Human Activity Recognition

Xue Wang, Yang Zhang

この論文をやさしく読む

ひとことで言うと

手首の動きから音が必要な場面だけマイクを短時間起動し、活動認識と会話の保護を両立しようとする方法。

何に役立つ?

ウェアラブル機器で音を補助情報に使う際、収音時間と発話の露出を減らす設計の参考になる。

この研究の面白いところ

マイクを常時使わず、異常検出で起動し、収録後にもマスクをかける二段階の構成。

どこまで分かった?

15人のデータによる管理されたオフライン評価で、実運用でのプライバシー保護はまだ検証されていない。発話認識に関する確認は小規模な予備試験である。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

音声には人の活動を示す豊富な手掛かりがあり、多くのウェアラブル機器にはすでにマイクが搭載されている。しかしマイクは会話も収録するため、プライバシー上の懸念が人間活動認識への利用を制限している。本研究は、手首のウェアラブル機器向けのセンサー起動法 AnomaSense を提示する。通常はマイクを切り、教師なし異常検出器が音を出しそうな慣性計測装置(IMU)の区間を検出したときだけ、最大1秒間マイクを入れる。収録した音は認識モデルへ送る前にさらにマスクする。15人の参加者による20種類の活動を調べた。活動は、手首の動きが似ているが使用する物体や材質が異なるものを5群に分けた。参加者を一人ずつ除外した検証では、IMU データだけの認識精度は78.98%だった。短くマスクした音を加えると、マスクしない場合は96.89%になり、各1秒の音の90%を除去しても86%を上回った。同じデータで、異常検出器が音の発生に対してマイクを起動する適合率は86.46%、再現率は74.28%だった。マスクした発話に対する自動音声認識についても小規模な予備確認を報告し、同じマスク率では、連続した区間を隠す方が点ごとに隠すより認識を大きく低下させた。評価は管理されたオフラインの実現可能性研究である。脅威モデル、手法が守れることと守れないこと、実環境での運用に必要な手順も説明する。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-24(UTC)
最新改訂
2026-09-24 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Audio carries rich cues about human activities, and microphones are already built into most wearable devices. However, microphones also capture speech, and this privacy risk limits their use in Human Activity Recognition (HAR). We present AnomaSense, a sensor activation approach for wrist wearables that keeps the microphone off by default and turns it on for at most one second when an unsupervised anomaly detector flags an IMU segment that is likely to produce sound. The captured audio is further masked before it reaches the recognition model. We study 20 activities from 15 participants, organized into five groups in which activities share similar wrist motion but differ in the object or material involved. With IMU data alone, our recognition model reaches 78.98% accuracy in leave-one-participant-out validation. With the short, masked audio windows added, accuracy reaches 96.89% with no masking and stays above 86% when 90% of each one-second audio window is removed. On the same data, the anomaly detector triggers the microphone with 86.46% precision and 74.28% recall relative to sound events. We also report a small preliminary check of automatic speech recognition on masked speech, which shows that contiguous masking degrades recognition far more than point-wise masking at the same masking ratio. Our evaluation is a controlled, offline feasibility study. We describe the threat model, what the approach does and does not protect, and the steps needed before deployment.

著者のコメント

12 pages, 6 figures

arXiv ID: 2609.28936 / 要約の誤りについて