1枚の画像から摩擦・硬さ・密度などを高速推定
PhysVGGT: Feed-Forward Dense Physical Property Estimation from A Single Image
この論文をやさしく読む
ひとことで言うと
写真1枚から、物体表面の摩擦や硬さなどの分布と、物体全体の質量を一度の計算で予測します。
何に役立つ?
ロボットが物体へ接触する前に、操作に必要な物性の見当をつける用途が考えられます。画像からの予測値であり直接測定ではありません。
この研究の面白いところ
局所的な物性を画素ごとに、質量を物体全体で推定し、疑似ラベルで訓練します。報告された推論時間は0.13秒で、従来比27倍の速度です。
どこまで分かった?
性能はABO-500とNeRF2Physicsで評価されています。要旨には物性ごとの誤差や実際のロボット操作の成功率はなく、画像から物性が一意に分かると証明したものではありません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
摩擦、硬さ、剛性、密度といった物理的性質は、ロボットが物体をどう把持し、操作し、相互作用すべきかを左右するが、それらをRGB画像から推定することは依然として難しい。既存手法は一般に、物理的性質を付加した物体ごとの再構成を用いるか、テスト時に視覚言語モデルへ直接問い合わせる。このため大きな計算負担が生じ、適用可能性が限られる。 本研究は、1枚のRGB画像から1回の順伝播で、摩擦係数、ショア硬さ、ヤング率、密度の密なマップと、物体単位の質量を予測するフィードフォワードモデルPhysVGGTを提示する。中心となる考えは、物理的性質の推定を画素ごとの密な予測問題として定式化することである。視覚幾何Transformerを使って入力画像から幾何を考慮したトークンを抽出し、その後、局所的な物理的性質を推定する密な予測分岐と、物体単位の質量を推定する大域的な予測分岐を用いる。さらに、密な物理的性質の予測を大規模な弱教師あり学習で訓練できる、拡張性のある疑似ラベル生成処理を導入し、高価な直接の物理測定への依存を大幅に減らす。広範な実験では、ABO-500データセットで最先端の性能を達成し、分布外のNeRF2Physicsデータセットにも有効に汎化することを示す。また、物体ごとの再構成とテスト時最適化を不要にし、画像1枚あたりわずか0.13秒の推論遅延を達成する。これは従来の最先端手法より27倍速い。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-16(UTC)
- 最新改訂
- 2026-09-16 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-16 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Physical properties, such as friction, hardness, stiffness, and density, govern how robots should grasp, manipulate and interact with objects, yet estimating these properties from RGB images remains challenging. Existing methods typically employ per-object reconstruction augmented with physical properties or directly query vision-language models at test time, which results in substantial computational overhead that limits their applicability. In this work, we present PhysVGGT, a feed-forward model that predicts dense maps of friction coefficient, Shore hardness, Young's modulus, and density, together with object-level mass, from a single RGB image in one forward pass. The key idea of PhysVGGT is to formulate physical property estimation as a dense per-pixel prediction problem and employ a visual geometry transformer to extract geometry-aware tokens from the input image followed by a dense prediction branch for estimating local physical properties and a global prediction branch for estimating object-level mass. In addition, we introduce a scalable pseudo-label generation pipeline that enables large-scale weakly supervised training for dense physical property prediction, substantially reducing the need for expensive direct physical measurements. Extensive experiments show that PhysVGGT achieves state-of-the-art performance on the ABO-500 dataset and generalizes effectively to the out-of-distribution NeRF2Physics dataset. Moreover, PhysVGGT eliminates the need for per-object reconstruction and test-time optimization, achieving an inference latency of only 0.13s per image, making it $27\times$ faster than the previous state of the art.
著者のコメント
Technical report
arXiv ID: 2609.18920 / 要約の誤りについて