幾何モデルを使う高精度な画像ベースのロボット位置合わせ
VGM-VS: Rethinking Visual Geometry Model for High-Precision Visual Servoing
この論文をやさしく読む
ひとことで言うと
事前学習した視覚幾何モデルとその場で集めた画像・姿勢データを組み合わせ、ロボットを画像から高精度に位置合わせする方法。
何に役立つ?
ケーブルや部品の把持・挿入など、高精度な位置合わせが必要な作業への適用が考えられる。実環境の3タスクで評価している。
この研究の面白いところ
視覚幾何モデルに残る並進距離の尺度の曖昧さを、ロボットが自分で集めた場面固有データで補正し、手先とカメラの関係も同時に学ぶ。
どこまで分かった?
精度と成功率は要旨に記載された3種類の組立タスクと試験条件での結果。ほかの作業や環境への一般化は要旨からは判断できない。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
本研究では、事前学習済みのフィードフォワード型視覚幾何モデルを使う画像ベースのロボット位置合わせ手法VGM-VSを提案する。現在の画像と目標位置で撮影した参照画像から視覚幾何モデルがカメラの相対姿勢を推定し、閉ループの姿勢ベース視覚サーボ(PBVS)でその推定値を位置姿勢の更新量として繰り返し適用する。大規模な事前学習で得た幾何を考慮した表現により、目標が隠れている、模様が少ない、画像中で小さいといった場合にも推定を保つ。 ただし、この種のモデルには尺度の曖昧さがあり、予測した並進は未知の倍率までしか定まらない一方、ロボット制御には実際の長さに基づく更新量が必要である。そこで、目標姿勢から始める所定の動きをロボットに自律実行させて画像と姿勢の組を記録し、その場面固有のデータでカメラ側のモデルを追加学習する。手先とカメラの変換も同時に学習するため、専用の較正工程を不要にする。 厳しい精度が必要な実環境の3タスク、USB-Cケーブルの把持、ケーブルの挿入、RAMの挿入で評価した。30 Hzでリアルタイム動作し、ケーブルのタスクでは最終位置が1 mm未満の精度に収束した。サーボ動作中に目標を動かした条件で成功率は90~100%だった。参照姿勢から最大30 cmずれた初期位置や、目標物の50%が隠れた条件では全試行で収束し、比較した視覚サーボ法より良い結果を示した。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-23(UTC)
- 最新改訂
- 2026-09-23 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-23 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
We present VGM-VS, a visual servoing method built on a pretrained feed-forward visual geometry model. Given the current view and a reference image captured at the target configuration, we estimate the relative camera pose with a visual geometry model and apply it iteratively as the pose increment of a closed-loop pose-based visual servoing (PBVS) scheme. The geometry-aware representation acquired from large-scale pretraining keeps this estimate reliable when the target is occluded, weakly textured, or covers only a small part of the image. However, the scale ambiguity inherent to these models leaves the predicted translation defined up to an unknown scale, while the pose increment must be metric for robot control. We close this gap with a scene-specific metric adaptation: the robot autonomously records image--pose pairs along a predefined motion starting from the target pose, and we fine-tune the camera head on these data, jointly learning the hand--eye transform and thus removing the need for a dedicated calibration process. We evaluate our method on three real-world assembly tasks with demanding tolerances: USB-C cable picking, cable insertion, and RAM insertion. Running in real time at 30Hz, VGM-VS converges to submillimeter terminal accuracy on the cable tasks, and reaches success rates of 90--100\% when the target is moved during servoing. It converges in all trials under initial displacements of up to 30cm from the reference pose and with 50\% of the target object occluded, outperforming the compared visual servoing baselines.
著者のコメント
8 pages, 3 figures. Corresponding author: Sen Wang
arXiv ID: 2609.28312 / 要約の誤りについて