自閉症の若者の行動が激しくなる前の焦燥を検出
Detecting Agitation Before Behavioral Escalation in Autistic Youth Through Multimodal Wearable Sensing
この論文をやさしく読む
ひとことで言うと
動き、生理情報、声をウェアラブル機器で測り、自閉症の若者が強い苦痛を感じている兆候を検出する研究です。15人のデータを1つの共有モデルで扱っています。
何に役立つ?
考えられる用途は、支援者が本人の苦痛の高まりに早く気付くための補助です。ただし、この研究で評価したのは検出性能であり、警告や介入によって行動の激化を防げたかではありません。
この研究の面白いところ
個人差の大きい状態に対して、事前学習モデルの特徴を使う共有モデルが有効だった点です。音声の寄与が大きく、腕時計だけでは同じ性能にならないことも比較しています。
どこまで分かった?
対象は15人・30セッションです。ROC曲線下面積は開始時0.724から30秒前0.608へ下がり、早い時点ほど判別が難しくなっています。これは正解率ではありません。日常環境での性能や支援効果は要旨には記載されていません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
攻撃、自傷、物の破壊などの困難を伴う行動は、自閉症の若者の68%に見られ、本人と介護者に危険をもたらす。これらのエピソードに先立って、動き、発声、自律神経の覚醒を通じて表れる、苦痛が高まった状態である焦燥が生じる。その兆候は微妙で個人差があり、自律神経の成分は計測機器なしでは見えない。 自閉症の若者15人を対象に、臨床家が主導する30セッションで、慣性計測装置から上半身の動き、手首装着機器から生理情報、襟元のマイクから発声を収集し、専門家による行動注釈と組み合わせた。各モダリティに1つずつ、4つの事前学習済み基盤モデルを適応させ、それぞれを共通の128次元空間に射影し、単一の集団モデルへ融合した。 このモデルは、臨床家が注釈した開始時点でROC曲線下面積0.724の焦燥検出性能を示し、参加者内の置換検定ではp=0.0005だった。開始30秒前では0.608へ低下した。15人中13人で偶然水準を上回った。ゼロから学習する構成は0.58にとどまった一方、特徴を固定した場合と微調整した場合の性能は同程度で、それぞれ0.71と0.72だった。信号への寄与は音声が最も大きく、腕時計だけの構成は偶然水準に近かった。 したがって、子どもごとに別のモデルを用意するのではなく、基盤モデルからの転移と単一の共有モデルを用いることで、注釈された開始時点より前の未注釈の時間窓を含め、個人ごとに異なる焦燥を検出できる。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-21(UTC)
- 最新改訂
- 2026-09-21 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-21 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Challenging behaviors including aggression, self-injury, and property destruction are observed in 68% of autistic youth and pose risks to youth and caregivers. These episodes are preceded by agitation, a rising state of distress expressed through movement, vocalization, and autonomic arousal. Its signs are subtle and individualized, and its autonomic components are invisible without instrumentation. We collected upper-body movement from inertial measurement units, physiology from a wrist-worn device, and vocalizations from lapel microphones across 30 clinician-led sessions with 15 autistic youth, paired with expert behavioral annotations. We adapt four pretrained foundation models, one per modality, project each to a shared 128-dimensional space, and fuse them into a single group model. The model detected agitation with an area under the ROC curve of 0.724 at the clinician-annotated onset (within-participant permutation p=0.0005), declining to 0.608 at 30,s before onset. Thirteen of fifteen participants were above chance. A from-scratch configuration reached only 0.58, while frozen and fine-tuned features performed comparably (0.71 and 0.72). Audio contributed most of the signal, and a watch-only configuration stayed near chance. Individualized agitation is therefore detectable, including in unannotated windows preceding the annotated onset, using foundation-model transfer with one shared model rather than one per child.
arXiv ID: 2609.24791 / 要約の誤りについて