arXiv論文メモ
新着一覧
cs.CV / cs.RO · 査読状況未確認

手の動きと物体の隠れ方から頭の動きを予測する

HAP: A Hand-Driven Active Perception Framework for Egocentric Head Motion Prediction

Yunji Feng, Junyi Ma, Guanzhong Sun, Chenyang Xu and Hesheng Wang

この論文をやさしく読む

ひとことで言うと

人が何に手を伸ばし、その対象がどう隠れるかを手掛かりに、頭が次にどちらへ動くかを予測します。頭の過去の動きだけに依存しない方法です。

何に役立つ?

考えられる用途は、人の操作行動に合わせた一人称システムの動作予測です。要旨で実証しているのはデータセット上の頭部運動の予測誤差の改善です。

この研究の面白いところ

対象の候補ごとに、現在と将来の隠れ方をグラフで表します。学習した予測と単純な等速度予測を、予測先の時間幅に応じて混ぜる点も特徴です。

どこまで分かった?

要旨には誤差の具体的な数値や改善率はありません。評価は公開データセットとBottleに基づき、あらゆる操作環境への適用性までは示していません。コードの公開は予定として記載されています。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

一人称視点の動作予測は主に手や操作対象の物体に注目してきた一方、将来の人間の頭部運動は比較的十分に研究されていない。物体を操作するとき、頭部は対象へ知覚を向け直して課題に必要な情報を得るとともに、身体や手の動きと協調する。そこで、観測した手の動きと推定した対象の状況を条件として、将来の頭部の6自由度(6-DoF)運動を予測する問題を定式化し、手の動きに基づく能動知覚の枠組みHAPを提案する。 HAPは、観測した手の動きと物体の幾何形状から、各物体が対象である確信度を推定する。次に、候補物体の間で現在生じている遮蔽と今後生じ得る遮蔽を表す、動的な予測対象中心アモーダル遮蔽グラフ(P-TAOG)を構成する。有向グラフと因果的な時間推論によって、対象に条件付けられて変化する知覚状態を符号化し、手と頭の運動履歴と融合する。さらに、予測の時間幅ごとのゲートを用いて、学習した軌道と等速度の事前仮定を混合する。 また、指定された対象に向けた物体操作を記録した一人称RGB-DデータセットBottleを導入する。これは、対象の見え方が変化する状況で協調する頭と手の動きを含む。公開データセットとBottleでの実験では、HAPが代表的な比較手法より小さい頭部運動予測誤差を達成した。この結果は、人間の頭部運動を予測するうえで、手の動きから意図を推定することと、動的な遮蔽を推論することの価値を支持する。コードは https://HAP-ego.github.io/HAP で公開予定である。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-16(UTC)
最新改訂
2026-09-16 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Egocentric motion forecasting has primarily focused on hands and manipulated objects, leaving future human head motion comparatively underexplored. During manipulation, the head both redirects perception toward the target to acquire task-relevant evidence and coordinates with body and hand motion. We therefore formulate future six Degree of Freedom (6-DoF) head-motion prediction conditioned on observed hand motion and inferred target context, and propose HAP, a Hand-Driven Active Perception framework. HAP infers confidence for each target object from observed hand motion and object geometry. Then constructs a dynamic Predictive Target-Centric Amodal Occlusion Graph (P-TAOG) representing current and potential occlusion among candidate objects. Directed graph and causal temporal reasoning encode the evolving target conditioned perceptual state, which is fused with hand and head motion history. A horizon-wise gate then blends the learned trajectory with a constant velocity prior. We further introduce Bottle, an egocentric RGB-D dataset of object manipulation toward specified targets, with coordinated head and hand motion under changing target visibility. Experiments on the public dataset and Bottle show that HAP achieves lower head motion prediction errors than representative baselines, supporting the value of hand driven intention and dynamic occlusion reasoning for anticipating human head motion. Code will be released at https://HAP-ego.github.io/HAP.

arXiv ID: 2609.18548 / 要約の誤りについて