arXiv論文メモ
新着一覧
cs.RO / cs.GR · 査読状況未確認

人の手とロボットの手で共有できる把持の幾何表現

InterMASH: A Unified Geometric Representation for Grasp Synthesis

Xuanze Yang and Yumeng Liu and Haiyang Xin and Changhao Li and Haowei Shen and Kai Xu and Ligang Liu and Ruizhen Hu

この論文をやさしく読む

ひとことで言うと

形が異なる人の手やロボットの手を、共通の幾何情報で表し、物体のつかみ方を生成する方法です。

何に役立つ?

考えられる用途は、人の把持データをロボットの把持学習へ活用することです。複数の手を一緒に学習できる表現を提案しています。

この研究の面白いところ

手の形と接触を別々に決めず、同じトークン表現で一緒に生成します。球面上のアンカーで、違う形の手の間に対応を作ります。

どこまで分かった?

要旨はShadowHandベンチマークでの評価を述べていますが、成功率の具体値や実機試験の条件は示していません。物理的妥当性の指標での評価と、任意の実物を確実につかめる保証は別です。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

把持生成は、安定し物理的に妥当な手と物体の相互作用を生成することを目指し、人の手のモデル化とロボット操作の両方で基本的な問題となっている。しかし、主に手の形態と表面モデルの違いから、人とロボットの手を統一して表す方法はまだ不足している。従来手法は、相互作用を表すために接触マップか密な陰的記述に頼ることが多いが、これらの表現は不完全だったり、計算が重く冗長だったりする。 球面上に固定したアンカーを使って、異なる身体構造の間の対応を作る統一的な幾何表現InterMASHを提案する。各アンカーでは、低次数の球面調和関数が、局所的な手の形状、物体の形状、接触をコンパクトに符号化し、明示的で解釈可能なトークン列を形成する。もともとトークン化されたこの構造に基づき、InterMASHの表現空間で直接動作する条件付きDiffusion Transformerを導入する。手の形状と接触を同時に生成し、整合性と物理的な妥当性を改善する。 大規模なShadowHandベンチマークの主要な物理的実現可能性指標で、最先端手法と競争力のある性能を達成し、複数種類の手を使う共同学習を可能にする。また、人の把持データで身体構造をまたいだ微調整を行うと、ロボットの把持成功と多様性を改善できることを示す。プロジェクトページはhttps://inter-mash.github.io/で公開されている。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-16(UTC)
最新改訂
2026-09-16 · v1
査読・掲載
査読状況未確認
arXivで読むPDFDOI

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Grasp synthesis aims to generate stable and physically plausible hand--object interactions, and has become a fundamental problem in both human hand modeling and robotic manipulation. However, a unified representation across human and robotic hands is still lacking, mainly due to differences in hand morphology and surface modeling. Prior methods typically rely on either contact maps or dense implicit descriptors to represent interaction, but these representations are often incomplete or computationally expensive and redundant. We propose InterMASH, a unified geometric representation that establishes cross-embodiment correspondence using sphere-fixed anchors. At each anchor, low-degree spherical harmonics compactly encode local hand geometry, object geometry, and contact, forming an explicit and interpretable token sequence. Building on this natively tokenized structure, we introduce a conditional Diffusion Transformer that operates directly in the proposed InterMASH representation space and jointly generates hand geometry and contact, improving consistency and physical plausibility. Our method achieves competitive performance with state-of-the-art methods on key physical feasibility metrics in a large-scale ShadowHand benchmark, supports joint training across multiple hands, and shows that cross-embodiment fine-tuning with human grasp data can improve robotic grasp success and diversity. Project page is available at https://inter-mash.github.io/.

著者のコメント

Project Page: https://inter-mash.github.io/

arXiv ID: 2609.18504 / 要約の誤りについて