arXiv論文メモ
新着一覧
cs.CV · 査読状況未確認

低解像度の歩行者画像から頭の向きを8分類する

Lightweight Pedestrian Head-Orientation Recognition Network for Safe Pedestrian-Vehicle Interaction

Yuanzhe Li, Yidi Huang, Xiaotong Chang, Hounian Liu

この論文をやさしく読む

ひとことで言うと

小さく写った歩行者の頭部画像から、顔や頭がどの方向を向いているかを8種類に分ける軽量な画像認識モデルです。

何に役立つ?

考えられる用途は、自動運転で歩行者の注意や行動を推定するときの入力情報です。実証の中心は頭の向きの分類であり、横断予測や事故低減そのものの成果とは分ける必要があります。

この研究の面白いところ

複数の公開データから頭部画像を集めて8方向のラベルを付け、低解像度向けモデルを比較しています。JAADとPIEでも交通場面での認識を評価しています。

どこまで分かった?

要旨には画像数、具体的な精度、モデルサイズ、処理速度は示されていません。最高精度という比較は、ここで評価したResNet-18、ResNet-34、VGG-16を含むモデル群の中での結果です。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

歩行者の頭の向きの認識は、歩行者の注意を理解し、横断する可能性を予測するための有用な手掛かりを与えるため、自動運転で重要な役割を持つ。しかし実際の交通場面では、歩行者の頭部領域が低解像度で撮影されることが多く、信頼できる認識は依然として難しい。 この課題に対処するため、歩行者の頭の向きを認識する軽量な低解像度頭部方向畳み込みニューラルネットワーク(LRHO-CNN)を提案する。複数の公開データセットから歩行者の頭部画像を抽出し、8つの方向カテゴリへ手作業で注釈を付けた新しいデータセットを構築する。収集画像には体系的な前処理と拡張を施し、データの多様性を増やすとともに、照明と画像品質の変動をよりよく表すようにする。 実験では、LRHO-CNNを、微調整したResNet-18、ResNet-34、VGG-16の三つの比較モデルと対比する。結果から、LRHO-CNNは評価したモデルの中で最も高い分類精度を達成する。さらにJAADとPIEデータセットでも評価し、実際の交通場面で歩行者の頭の向きを認識する有効性と、後段の歩行者行動・意図予測を支えうる有用な頭部方向の手掛かりを提供することを示す。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-21(UTC)
最新改訂
2026-09-21 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Pedestrian head orientation recognition plays an important role in autonomous driving by providing valuable cues for understanding pedestrian attention and anticipating potential crossing behavior. However, reliable recognition in real-world traffic scenes remains challenging because pedestrian head regions are often captured at low resolution. To address this challenge, we propose a lightweight Low-Resolution Head Orientation Convolutional Neural Network (LRHO-CNN) for pedestrian head orientation recognition. We construct a new dataset by extracting pedestrian head images from multiple public datasets and manually annotating them into eight orientation categories. The collected images are systematically preprocessed and augmented to increase data diversity and better represent variations in illumination and image quality. The experimental analysis compares LRHO-CNN with three fine-tuned baseline models, namely ResNet-18, ResNet-34, and VGG-16. The results demonstrate that LRHO-CNN achieves the highest classification accuracy among the evaluated models. LRHO-CNN is further evaluated on the JAAD and PIE datasets, demonstrating its effectiveness in recognizing pedestrian head orientation in real-world traffic scenes and providing informative head-orientation cues that can support downstream pedestrian behavior and intention prediction.

arXiv ID: 2609.24193 / 要約の誤りについて