arXiv論文メモ
新着一覧
cs.CV · 査読状況未確認

姿勢と画像を組み合わせ歩行者の携帯電話による注意散漫を検出

Detecting Phone-Induced Pedestrian Distraction via a Multimodal Fusion Transformer

Yuanzhe Li, Hounian Liu, Xiaotong Chang, Yidi Huang

この論文をやさしく読む

ひとことで言うと

歩行者の姿勢の変化と見た目の画像を組み合わせて、携帯電話の使用による注意散漫を判別する方法です。両情報の関係と時間変化をTransformerで扱います。

何に役立つ?

想定用途は、自動運転車が歩行者の状態を把握し、リスク評価に使うことです。実証されているのは注釈付きデータセットでの分類性能で、事故の減少を直接測った結果ではありません。

この研究の面白いところ

姿勢だけ、外観だけで判断せず、交差注意で補い合う情報を取り込みます。その後に時間的な依存を捉える構成になっています。

どこまで分かった?

95%は287事例・20,741画像のデータセットでの全体正解率です。原文の「6%上回る」は相対改善率かパーセントポイント差か明記されていないため、換算していません。実路での安全性や別環境への一般化は要旨に示されていません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

携帯電話への依存が増すにつれて、電話の使用による歩行者の注意散漫がますます一般的になっている。メッセージ入力、動画視聴、通話などの行為は、交通事故の重要な要因となっている。歩行者の注意散漫を確実に検出することは、自動運転車にとって不可欠である。状況認識を高め、適時のリスク評価を可能にし、安全な動作計画と車両制御を支えるためである。 携帯電話による歩行者の注意散漫を検出する、マルチモーダル融合Transformer(MFT)を提案する。MFTは、身体姿勢のキーポイントから骨格の動態を、歩行者画像から外観特徴を同時に抽出し、二つのモダリティが持つ相補的な情報を活用する。マルチヘッド交差注意によってモダリティ間の依存を捉える交差モダリティ注意モジュールを提案し、両者の相補的な情報を効果的に融合する。さらに、Transformerエンコーダで実装した時間的注意融合モジュールを用い、時間に沿った依存を捉える。 MFTは、手作業で注釈を付けた287の歩行者事例、20,741枚の画像からなるデータセットで学習・評価した。広範な実験で、全体の正解率は95%に達し、六つの基準手法の性能を6%上回った。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-20(UTC)
最新改訂
2026-09-20 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

The increasing reliance on mobile phones has made phone-induced pedestrian distraction increasingly prevalent. Activities such as texting, watching videos, and making phone calls have become significant contributors to traffic accidents. Reliable detection of pedestrian distraction is essential for autonomous vehicles, as it improves situational awareness and enables timely risk assessment, thereby supporting safe motion planning and vehicle control. We propose a multimodal fusion Transformer (MFT) for detecting phone-induced pedestrian distraction. MFT jointly extracts skeletal dynamics from body pose keypoints and visual appearance features from pedestrian images, effectively leveraging the complementary information provided by the two modalities. A cross-modal attention module is proposed to capture inter-modal dependencies through multi-head cross-attention, facilitating effective fusion of complementary information across the two modalities. Then, a temporal attention fusion module, implemented with a Transformer encoder, is employed to capture temporal dependencies. MFT is trained and evaluated on a manually annotated dataset comprising 287 pedestrian instances with 20,741 images. Extensive experiments demonstrate that MFT attains an overall accuracy of 95%, exceeding the performance of six baseline approaches by 6%.

著者のコメント

2026 IEEE 9th International Conference on Electronic Information and Communication Technology (ICEICT 2026)

arXiv ID: 2609.23507 / 要約の誤りについて