単眼画像から手首のカメラ座標を含む手の姿勢を推定
Estimating Accurate Hand Pose in Camera Space with Vision Transformer
この論文をやさしく読む
ひとことで言うと
一台のカメラの画像から、手の形とカメラに対する手首の位置を併せて推定する方法。
何に役立つ?
考えられる用途はカメラ座標系で手の動きを扱う視覚システム。要旨ではHO3Dでの評価を示している。
この研究の面白いところ
奥行きの曖昧さと、透視投影で局所姿勢と手首位置が絡む問題に別々の仕組みを設ける。
どこまで分かった?
示される最大37.1%の改善はHO3DのCS-MJE指標での比較。ほかの条件で同じ改善率になるとは要旨からは分からない。
v2のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
単眼RGB画像に基づく手の姿勢推定は、コンピュータービジョンで重要な研究分野になっている。局所的な手の姿勢推定では手首を基準に姿勢を予測するが、大域的な推定ではカメラ座標系での手首の位置も求める必要がある。しかし、カメラ空間での推定には、単眼画像での奥行きの曖昧さと、透視投影で局所的な手の姿勢と手首の大域位置が結び付くという二つの課題がある。この結び付きは、投影像が局所姿勢、手首位置、カメラ内部パラメータで共同して決まることを意味する。 これらに対処するため、著者らは手の奥行き情報を抽出するTransformation-Isomorphism Supervisionと、局所姿勢と手首位置の結び付きを扱うPerspective Information Embeddingを提案し、一般的なエンコーダー・デコーダー構成に組み込む。さらに、連続画像での姿勢を改善するため、フレームレートを考慮した複数データセットの新しい学習方法を提案する。統合した手法は、HO3DでのCS-MJE指標において、従来の最先端手法より最大37.1%優れた結果を得た。プロジェクトページは要旨記載のGitHubにある。
v2の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-21(UTC)
- 最新改訂
- 2026-09-22 · v2
- 査読・掲載
- 査読状況未確認
更新履歴
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Monocular RGB-based hand pose estimation has emerged as a critical research frontier in computer vision. The local hand pose estimation methods predict hand poses relative to the wrist, while global hand pose estimation also requires estimating the wrist's position in the camera coordinate system. However, this camera-space estimation confronts two fundamental challenges: (1) depth ambiguity in monocular settings, and (2) the coupling effect of hand local poses and global wrist positions in the perspective projections. In particular, this coupling reflects that the projections are jointly determined by local hand poses, wrist positions, and camera intrinsics. To overcome these challenges, our framework proposes two key innovations: Transformation-Isomorphism Supervision for hand-depth information extraction and Perspective Information Embedding for resolving above coupling effect of local pose and wrist position, both integrated within the mainstream encoder-decoder architecture. Besides, we propose a novel framerate-aware multi-dataset training strategy for sequential pose refinement. Our fully integrated approach achieves at most 37.1\% superiority in CS-MJE over SOTA on HO3D. Project page: https://github.com/Mine268/CS-ViT.
arXiv ID: 2609.24424 / 要約の誤りについて