人が触って測った力からロボット用の物体モデルを作る
ForceTwin: Physics-informed Digital Twins for Robotic Manipulation from Instrumented Human Interaction
この論文をやさしく読む
ひとことで言うと
人が力を測りながら扉などを動かし、見た目だけでは分からない重さや摩擦、ばねの力をロボット用の仮想モデルに取り込む研究です。
何に役立つ?
抵抗の強い扉などをロボットが操作するために、物体の動きと必要な力を予測するモデルとして役立ちます。要旨では実機制御と、扉通過方策の実環境への展開を報告しています。
この研究の面白いところ
固定した物理パラメータに加え、姿勢や速度で変わる機構力をニューラル残差として表します。画像からの推定だけでは見落とす抵抗を、人の計測動作から取り込む点が特徴です。
どこまで分かった?
87%の目標達成率は9通りの物体・機体の組合せに関する値です。慣性パラメータ誤差のほぼ半減と、操作の成功率は異なる評価です。要旨には各組合せの試行数や、対象外の物体への性能は示されていません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
物体を操作するには、その動きだけでなく、動きを決める物理的特性も理解する必要がある。関節を持つ物体では、慣性、摩擦、ばねやドアクローザーなどの機構がこれに含まれ、その作用は姿勢や速度によって変化し得る。こうした特性は外見から直接観察できず、見た目が同じ扉でも、操作に必要な力は大きく異なる場合がある。既存のデジタルツイン構築手法は、主に運動学を復元するか、視覚・言語の事前情報から固定の物理パラメータを割り当てるため、物理的に不自然な推定になることがある。その結果、状態に依存する機構のダイナミクスは同定されず、標準的なアセット形式にも表現されない。 本研究では、計測器を使った人間の相互作用から、関節を持つ物体の物理情報を組み込んだデジタルツインを同定するシステムForceTwinを提案する。人が手持ちの力センサ付きグリッパーで物体を動かして調べることで、同期した位置・姿勢と相互作用力を得る。そこから、関節構造、慣性・クーロン摩擦・粘性減衰を含むパラメトリックなダイナミクス、さらに状態依存の機構力を捉える構造化されたニューラル残差を推定する。ForceTwinは、VLMによる事前推定と比べ、慣性パラメータの誤差をほぼ半減させる。 SpotおよびFranka FR3のインピーダンス制御におけるフィードフォワードの動力学モデルとして利用したところ、9通りの物体とロボット機体の組合せで、ForceTwinの目標達成率は87%だった。VLMの事前推定を使う場合は60%、運動学のみのツインを使う場合は57%であり、機構の作用が強いため両ベースラインが途中で止まる物体で、改善が最も大きかった。さらに、同定したツインを使って全身を用いる扉通過の方策を学習し、実環境へ展開した。プロジェクトページは https://timengelbracht.github.io/forcetwin-website/ である。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-18(UTC)
- 最新改訂
- 2026-09-18 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-18 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Manipulating objects requires understanding not only their motion, but also the physical properties that determine it. For articulated objects, these include inertia, friction, and mechanisms such as springs or door closers, whose effects can vary with configuration and velocity. Such properties are not directly observable from appearance: visually identical doors may require very different effort to manipulate. Existing digital-twin pipelines recover primarily kinematics or assign static physical parameters from visual and language priors, which can yield physically implausible estimates. As a result, state-dependent mechanism dynamics remain unidentified and are not represented in standard asset formats. We present ForceTwin, a system for identifying physics-informed digital twins of articulated objects from instrumented human interaction. A person probes an object using a handheld force-sensing gripper, providing synchronized poses and interaction forces from which we estimate the articulation, parametric dynamics including inertia, Coulomb friction, viscous damping, and a structured neural residual capturing state-dependent mechanism forces. ForceTwin nearly halves the inertial-parameter error of a VLM prior. As a feedforward dynamics model for impedance control on a Spot and a Franka FR3, ForceTwin achieves 87% goal completion across nine object-embodiment pairs, compared with 60% using VLM-prior and 57% using kinematics-only twins, with the largest gains on objects whose strong mechanisms cause both baselines to stall. We further use the identified twins to train whole-body door-traversal policies and deploy them in the real world. Project Page: https://timengelbracht.github.io/forcetwin-website/
arXiv ID: 2609.21751 / 要約の誤りについて