日常作業の一人称映像から両腕ロボットの操作を学ぶ
EgoWild2Dex: Learning Dexterous Robotic Manipulation from In-the-Wild Human Experience
この論文をやさしく読む
ひとことで言うと
日常の作業映像をロボットの見え方と動きへ合わせ、両腕と器用な手による操作の学習に使う方法です。
何に役立つ?
考えられる用途は、人間の既存の作業経験を使ってロボット実演データの必要量を抑えることです。要旨では実ロボット3課題で評価しています。
この研究の面白いところ
頭の動きによる視点の揺れと、人間・ロボットの身体の違いを同時に扱い、大規模な日常映像を利用します。
どこまで分かった?
96.7%は3課題の平均で、未知物体に対するゼロショット成績は33.3%です。未知の物体すべてで高成功率を示したわけではなく、末尾ではデータ等は公開予定とされています。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
人間の一人称視点データは、器用なロボット操作を学習するための理にかなった教師情報となる。従来の手法は、制約された環境や特別に構築した環境でデータを集めることが多かった。これに対し本研究では、家庭、工場、薬局などの現実の場で、人が頭部装着カメラを着けて通常の仕事を行う、日常環境の一人称デモンストレーションを収集する。この収集方法は、多様な作業手順と手・物体の相互作用を、出現頻度の低い物体や技能まで含む分布にわたって捉える。一方、場面の混雑や頭の動きによる視点変化により、視覚的に難しい観測にもなる。平均累積回転量は毎秒15.93度である。 これらの問題に対処するため、EgoWild2Dexを導入する。これは、不安定な人間の一人称視点をロボットの観測へ、人間の動作をロボットの行動へ、それぞれ同時に整合させることで、日常環境の人間の経験を器用な手を持つ双腕ロボットへ移す。本研究は3つの利点を提供する。第1に、雑音を含む人間の観測をロボットの観測に近づけて変形する、微分可能な幾何TransformerであるGeoFormerを導入する。第2に、身体構造の違いを埋め、少量のロボット側の教師情報で高い課題成功率を可能にする、人間・ロボットの訓練方式を設計する。第3に、179,049エピソード、125,961種類の課題記述、1,282種類の物体カテゴリからなる、538.9時間の日常環境の一人称人間データセットEgoWildを公開する。 実ロボットでは、EgoWild2Dexは長期にわたる両手の器用な操作3課題で平均成功率96.7%を達成し、物体単位のゼロショット平均成功率は33.3%だった。データ、モデル、コードは公開予定である。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-20(UTC)
- 最新改訂
- 2026-09-20 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-20 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Egocentric human data provide a principled source of supervision for learning dexterous robot manipulation. Unlike prior approaches that often collect such data in constrained or specially constructed environments, we collect in-the-wild egocentric demonstrations in real-world settings, including homes, factories, and pharmacies, etc., where people perform their ordinary tasks while wearing head-mounted cameras. This collection protocol captures diverse workflows and hand-object interactions across long-tailed object and skill distributions, but also yields visually challenging observations due to scene clutter and head-motion-induced viewpoint changes (a mean cumulative rotation of $15.93^{\circ}$/s). To address these issues, we introduce EgoWild2Dex, which transfers in-the-wild ego-human experience to dual-arm robots with dexterous hands by jointly aligning unstable egocentric views and human motions with robot observations and actions, respectively. This work offers three benefits. First, we introduce GeoFormer, a differentiable geometric transformer that warps noisy human observations toward robot observations. Second, we design a human-robot training scheme to bridge the embodiment gap, enabling high task success with limited robot supervision. Third, we release EgoWild, a 538.9-hour in-the-wild egocentric human dataset comprising 179,049 episodes, 125,961 unique task descriptions, and 1,282 object categories. On real robots, EgoWild2Dex achieves an average success rate of 96.7% across three long-horizon bimanual dexterous manipulation tasks and an average object-level zero-shot success rate of 33.3%. The data, models, and code will be released.
arXiv ID: 2609.23755 / 要約の誤りについて