日常動画から可動物体と手の動きを復元する
Track, Articulate, Act: Generating Articulation from Casual Human Videos
この論文をやさしく読む
ひとことで言うと
ドアや引き出しを人が動かす普通の動画から、関節付き3次元モデルと手の動きを復元する方法です。
何に役立つ?
実証されているのは復元モデルと手の軌跡を使ったMuJoCo内での相互作用の再生です。考えられる用途は、ロボット操作研究に使うシミュレーション素材の作成です。
この研究の面白いところ
固定部分と動く部分の点の軌跡を手がかりに、既存の視覚モデルの予測を幾何学でつなぎ、関節を人手で指定せずに推定する点です。
どこまで分かった?
要旨には復元精度の数値や成功率はありません。シミュレーションでの再生と、実機ロボットが同じ操作を達成することは区別する必要があります。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
人の動画には、ロボット操作のための豊かな因果的情報が含まれる。手の動きが物体の動きをどう引き起こし、課題に関係する物体状態の変化をどう生むかを示すからである。本研究では、ドア、引き出し、戸棚、ノートパソコン、オーブン、蝶番付き容器など、日常に広く存在し、身体を介した相互作用に固有の課題をもたらす関節物体を扱う。これらの物体は単一の姿勢では表せず、その動きは内部の部品と関節に依存する。 日常的に撮影された単眼RGB動画から、シミュレーションで利用可能な関節物体と手・物体間の相互作用を復元する、実世界からシミュレーションへの枠組みを導入する。RGB-Dや多視点入力、事前のスキャン、手動で指定した関節、ロボットによる実演は必要ない。重要な着想は、密な3次元点の軌跡が、身体の形態に依存しない関節構造の手がかりになることである。固定部品上の点はほぼ静止し、可動部品上の点は一貫した回転運動または直動運動をする。本手法は部品を分割し、関節とその状態軌跡を推定し、関節付きの物体モデルを再構成して、復元した3次元の手の動きを物体に整合させる。 中心となるのは、単一画像からの3次元復元、メッシュ分割、3次元シーンフローの強力な事前学習モデルを転用するモジュール型の構成である。これらの予測を明示的な幾何学的推論で結び付け、関節構造を推定する。復元した物体と人の手の軌跡を用い、MuJoCo内で接触を通じて相互作用を再生する。この枠組みは、事前学習済み視覚モデルと明示的な運動推論により、日常動画を、その後の身体を介した相互作用に適した関節物体モデルへ変換できることを示す。https://track-articulate-act.github.io/
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-16(UTC)
- 最新改訂
- 2026-09-16 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-16 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Human videos contain rich causal evidence for robot manipulation: they reveal how hand motion induces object motion and produces task-relevant changes in object state. In this work, we study articulated objects such as doors, drawers, cabinets, laptops, ovens, and hinged containers that are ubiquitous in daily life and present unique challenges for embodied interaction. These objects cannot be represented by a single pose; their motion depends on the underlying parts and joints. We introduce a real-to-sim framework that reconstructs a simulation-ready articulated object and hand-object interaction from a casual monocular RGB video, without RGB-D or multi-view input, prior scans, manually specified joints, or robot demonstrations. Our key insight is that dense 3D point tracks provide an embodiment-agnostic articulation cue: points on the fixed link remain approximately stationary, while points on the moving link follow coherent revolute or prismatic motion. Our method segments the links, estimates the joint and its state trajectory, reconstructs an articulated asset, and aligns the recovered 3D hand motion with the object. Central to our approach is a modular recipe that repurposes powerful pretrained models for single-image 3D reconstruction, mesh segmentation, and 3D scene flow, connecting their predictions through explicit geometric reasoning to infer articulation. We use the reconstructed articulated object and the human hand trajectory to replay interactions through contact in MuJoCo. The framework shows how pretrained vision models and explicit motion reasoning can turn casual human videos into articulated object models suitable for downstream embodied interactions. https://track-articulate-act.github.io/
著者のコメント
Preprint. Under Review
arXiv ID: 2609.19119 / 要約の誤りについて