arXiv論文メモ
新着一覧
cs.RO · 査読状況未確認

人の動画からロボット操作への転移を比べるH2RBench

H2RBench: A Real-to-Sim Benchmark for Evaluating Human-to-Robot Transfer

Chuyang Xiao, Haotian Zhan, Sriram Krishna, Peilin Meng, Muhammad Zubair Irshad, Sergey Zakharov, David Held

この論文をやさしく読む

ひとことで言うと

人の作業動画からロボットへ操作を移す手法を、共通条件で比較するためのベンチマークです。

何に役立つ?

転移手法を実ロボットで試す前に、同じ課題とデータ量で比較するのに役立ちます。

この研究の面白いところ

人の実演を増やしたときの効果が手法間で大きく違い、シミュレーション成績と実ロボット成績には高い相関がありました。

どこまで分かった?

課題は四つの操作課題です。相関は評価した手法・課題の組み合わせでの値で、あらゆるロボット作業への予測力を保証しません。

v2のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

人の実演動画からロボットの操作方策を学ぶことは、ロボット学習を大規模化する有望な道筋である。しかし、人からロボットへの転移手法は、課題群、場面の配置、物体の種類、ロボット側の教師情報の量など異なる条件で評価されており、比較が難しい。この問題に対処するため、実環境からシミュレーションへの変換を使い、人からロボットへの転移を評価するH2RBenchを提示する。実際の人の実演動画とシミュレーション上のロボット実演に基づく標準的な手順を提供し、異なる相互作用を要する四つの操作課題を含む。身体構造の違いを埋める方法が異なる複数の代表的な転移手法を評価した。H2RBenchにより、人の実演数を増やしたときの性能変化を体系的に調べ、追加の人のデータを活用する能力が手法ごとに大きく異なることを示した。さらに、シミュレーションでの性能は実ロボットでの性能をおおむね予測し、手法と課題の組み合わせ全体でPearsonの相関係数は0.89、Spearmanの相関係数は0.85、平均最大順位違反MMRVは0.06だった。これらの結果は、実環境への導入前に転移手法を比較するため、H2RBenchが実用的で規模を拡張しやすいベンチマークであることを示す。

v2の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-21(UTC)
最新改訂
2026-09-22 · v2
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Learning robot manipulation policies from human video demonstrations constitutes a promising avenue for scalable robot learning. However, comparing different human-to-robot (H2R) transfer methods remains challenging, as existing approaches are evaluated under different settings, including differing task suites, scene layouts, object instances, and amounts of robot supervision. To address this challenge, we present H2RBench, a Real2Sim benchmark for evaluating H2R transfer methods. H2RBench provides a standardized protocol built on real human video demonstrations and simulated robot demonstrations, and includes four manipulation tasks spanning diverse interaction requirements. We evaluate multiple representative H2R transfer methods, each adopting a different strategy for bridging the embodiment gap. Using H2RBench, we systematically characterize how each method scales with the amount of human demonstrations, revealing that methods differ substantially in their ability to leverage additional human data. We further show that simulation performance is broadly predictive of real-world robot performance, with an overall Pearson correlation of r = 0.89, Spearman correlation of \r{ho} = 0.85 and Mean Maximum Rank Violation (MMRV) of 0.06 across method-task configurations. These results establish H2RBench as a practical and scalable benchmark for comparative H2R evaluation prior to real-world deployment.

著者のコメント

10th Conference on Robot Learning (CoRL 2026), Austin, TX, USA

arXiv ID: 2609.24778 / 要約の誤りについて