arXiv論文メモ
新着一覧
cs.LG · 査読状況未確認

装着型センサーによる感情認識のモデルと測定条件を比較

From Stress to Affect: Multimodal Deep Learning for Physiological Emotion Recognition Across Wearable Sensor Modalities

Desta Haileselassie Hagos, Saurav Keshari Aryal, Legand L. Burge

この論文をやさしく読む

ひとことで言うと

手首や胸部の生理信号から感情を分類するとき、どのモデルやセンサー構成がよいかを二つのデータセットで比較しています。

何に役立つ?

感情認識システムで、測定する信号と計算コストのバランスを考える参考になります。特に4 Hzという低い測定頻度での性能を検討しています。

この研究の面白いところ

Transformerが常に最良とは限らず、データセットによってLSTMが上回ることを、参加者を分けた評価で示しています。

どこまで分かった?

結果は二つのデータセットに対する交差検証です。高い分類正解率はメンタルヘルスの診断精度や日常環境での有効性を直接示すものではありません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

装着型センサーを用いた生理信号による感情認識は、メンタルヘルスのモニタリング、感情コンピューティング、人とコンピューターの相互作用に重要な用途を持つ。しかし既存研究は、通常、単一のモデル、センシング構成、またはデータセットを評価しており、これらの要因が認識性能にどう影響するかについての理解は限られている。本研究では、二つの複数モダリティの装着型センサーデータセットWESADとEmoWearを用い、生理信号による感情認識のための時系列深層学習構造を比較する。 双方向の長短期記憶(LSTM)、時間畳み込みネットワーク(TCN)、Transformerを、手首のみ、胸部のみ、複数モダリティというセンシング構成で評価した。評価には、参加者が重ならない一人抜き交差検証(LOSO-CV)を用いた。さらに、ソフト投票によるアンサンブル、センサーの除去、サンプリング周波数、勾配に基づくサリエンシーを調べた。 WESADではTransformerが複数モダリティで最も高い正解率99.02%±0.51%を達成した。一方、EmoWearではLSTMが複数モダリティで、覚醒度91.80%±1.06%、感情価89.96%±0.36%の双方について最良の正解率を達成した。これらの結果は、一つの構造が一律に優れるのではなく、相対的な性能がデータセットの特徴に依存することを示している。複数モダリティのセンシングは、両データセットで手首のみと胸部のみの構成を一貫して上回った。サンプリング周波数の分析から、4 Hzは、より高い周波数と同程度の性能を大幅に低い訓練コストで得られる、実用的な動作点であることが分かった。本知見は、装着型センサーによる生理学的な感情認識において、モデル構造、センシングモダリティ、サンプリング周波数を選ぶための指針を与える。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-17(UTC)
最新改訂
2026-09-17 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Physiological emotion recognition using wearable sensors has important applications in mental health monitoring, affective computing, and human-computer interaction. However, existing studies typically evaluate a single model, sensing configuration, or dataset, limiting our understanding of how these factors influence recognition performance. We present a comparative study of temporal deep learning architectures for physiological emotion recognition using two multimodal wearable datasets: WESAD and EmoWear. Bidirectional long short-term memory (LSTM), temporal convolutional network (TCN), and Transformer models are evaluated under wrist-only, chest-only, and multimodal sensing configurations using participant-independent leave-one-subject-out cross-validation (LOSO-CV). We also investigate soft-voting ensembles, sensor ablation, sampling frequency, and gradient-based saliency. The Transformer achieved the highest multimodal accuracy on WESAD (99.02% +/- 0.51%), whereas the LSTM achieved the best multimodal accuracy on EmoWear for both arousal (91.80% +/- 1.06%) and valence (89.96% +/- 0.36%). These results show that relative architecture performance depends on dataset characteristics rather than one architecture being uniformly superior. Multimodal sensing consistently outperformed wrist-only and chest-only configurations across both datasets. Sampling-frequency analysis showed that 4 Hz provides a practical operating point, with performance comparable to higher frequencies at substantially lower training cost. These findings provide guidance for selecting architectures, sensing modalities, and sampling frequencies for wearable physiological emotion recognition.

著者のコメント

Under review

arXiv ID: 2609.20991 / 要約の誤りについて