arXiv論文メモ
新着一覧
cs.HC / cs.LG · 査読状況未確認

自閉症の若者の行動が激しくなる前の焦燥を検出

Detecting Agitation Before Behavioral Escalation in Autistic Youth Through Multimodal Wearable Sensing

Nibraas Khan, Abigale Plunk, John Staubitz, Ingrid Shragge, Jordan Brooks, Suzanne Wright, Alec Brewer, James Dieffenderfer, Alper Bozkurt, Amy Weitlauf, Nilanjan Sarkar

この論文をやさしく読む

ひとことで言うと

動き、生理情報、声をウェアラブル機器で測り、自閉症の若者が強い苦痛を感じている兆候を検出する研究です。15人のデータを1つの共有モデルで扱っています。

何に役立つ?

考えられる用途は、支援者が本人の苦痛の高まりに早く気付くための補助です。ただし、この研究で評価したのは検出性能であり、警告や介入によって行動の激化を防げたかではありません。

この研究の面白いところ

個人差の大きい状態に対して、事前学習モデルの特徴を使う共有モデルが有効だった点です。音声の寄与が大きく、腕時計だけでは同じ性能にならないことも比較しています。

どこまで分かった?

対象は15人・30セッションです。ROC曲線下面積は開始時0.724から30秒前0.608へ下がり、早い時点ほど判別が難しくなっています。これは正解率ではありません。日常環境での性能や支援効果は要旨には記載されていません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

攻撃、自傷、物の破壊などの困難を伴う行動は、自閉症の若者の68%に見られ、本人と介護者に危険をもたらす。これらのエピソードに先立って、動き、発声、自律神経の覚醒を通じて表れる、苦痛が高まった状態である焦燥が生じる。その兆候は微妙で個人差があり、自律神経の成分は計測機器なしでは見えない。 自閉症の若者15人を対象に、臨床家が主導する30セッションで、慣性計測装置から上半身の動き、手首装着機器から生理情報、襟元のマイクから発声を収集し、専門家による行動注釈と組み合わせた。各モダリティに1つずつ、4つの事前学習済み基盤モデルを適応させ、それぞれを共通の128次元空間に射影し、単一の集団モデルへ融合した。 このモデルは、臨床家が注釈した開始時点でROC曲線下面積0.724の焦燥検出性能を示し、参加者内の置換検定ではp=0.0005だった。開始30秒前では0.608へ低下した。15人中13人で偶然水準を上回った。ゼロから学習する構成は0.58にとどまった一方、特徴を固定した場合と微調整した場合の性能は同程度で、それぞれ0.71と0.72だった。信号への寄与は音声が最も大きく、腕時計だけの構成は偶然水準に近かった。 したがって、子どもごとに別のモデルを用意するのではなく、基盤モデルからの転移と単一の共有モデルを用いることで、注釈された開始時点より前の未注釈の時間窓を含め、個人ごとに異なる焦燥を検出できる。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-21(UTC)
最新改訂
2026-09-21 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Challenging behaviors including aggression, self-injury, and property destruction are observed in 68% of autistic youth and pose risks to youth and caregivers. These episodes are preceded by agitation, a rising state of distress expressed through movement, vocalization, and autonomic arousal. Its signs are subtle and individualized, and its autonomic components are invisible without instrumentation. We collected upper-body movement from inertial measurement units, physiology from a wrist-worn device, and vocalizations from lapel microphones across 30 clinician-led sessions with 15 autistic youth, paired with expert behavioral annotations. We adapt four pretrained foundation models, one per modality, project each to a shared 128-dimensional space, and fuse them into a single group model. The model detected agitation with an area under the ROC curve of 0.724 at the clinician-annotated onset (within-participant permutation p=0.0005), declining to 0.608 at 30,s before onset. Thirteen of fifteen participants were above chance. A from-scratch configuration reached only 0.58, while frozen and fine-tuned features performed comparably (0.71 and 0.72). Audio contributed most of the signal, and a watch-only configuration stayed near chance. Individualized agitation is therefore detectable, including in unannotated windows preceding the annotated onset, using foundation-model transfer with one shared model rather than one per child.

arXiv ID: 2609.24791 / 要約の誤りについて