arXiv論文メモ
新着一覧
cs.RO · 査読状況未確認

物体の意味と位置を分けてロボット操作を安定化

Object-Centric Conditioning for Visuomotor Flow Matching

Jijie Li, Xu Yang, Junhong Zou, Chunhai Zhao, Chaoyang Zhao, Zhen Lei, and Xiangyu Zhu

この論文をやさしく読む

ひとことで言うと

ロボットが過去の動きを再利用するとき、目の前の物体が何でどこにあるかを別々に捉え、配置の変化に対応させます。

何に役立つ?

背景に邪魔な物があったり対象の位置がずれたりする操作で、少ない推論ステップを維持しつつ安定性を高める用途があります。

この研究の面白いところ

過去の動作を捨てず、現在の物体情報で補い、どの構成要素が改善に効いたかも切り分けています。

どこまで分かった?

シミュレーションと実世界の両方で検証していますが、要旨には試行数や成功率の数値はありません。任意の環境変化への保証ではありません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

ロボットの視覚運動方策は、自己回帰モデル、拡散モデル、あるいは近年ではフローマッチングモデルとして定式化されることが多い。その中でAction-to-Action(A2A)フローマッチングは、確率的な雑音ではなく、過去の行動の事前情報から生成を始めることで推論効率を改善する。しかし、古くなった過去の動作パターンと、全体の視覚情報が絡み合った表現は、空間的な分布外(OOD)の変化や視覚的な妨害要素の下で、頑健性をともに低下させる可能性がある。 本研究では、頑健な視覚運動操作のための、物体中心のフローマッチング方策SlotFlowを提案する。SlotFlowは、場面の観測を意味的な「何であるか」の特徴と、画像平面上の軽量な「どこにあるか」の位置手掛かりへ分離する。これにより、物体を考慮した方策の条件付けと、現在状態への対応付けを行う。意味表現は無関係な背景との相関を抑え、位置手掛かりは物体配置の変化への適応を改善する。広範なシミュレーションと実世界実験により、A2Aの少ないステップでの推論効率を保ちながら、視覚的妨害や大きな位置の変化に対する頑健性が改善することを示す。初期化と知覚を統制して構成要素を取り除く実験から、物体中心の現在状態への対応付けが改善の主要因であり、有用な過去の動作の事前情報を置き換えるのではなく補完することも示す。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-21(UTC)
最新改訂
2026-09-21 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Robot visuomotor policies are commonly formulated as autoregressive, diffusion-based, or more recently, flow matching models. Among them, Action-to-Action (A2A) flow matching improves inference efficiency by initializing generation from historical action priors rather than stochastic noise. However, stale historical motion patterns and entangled global visual representations can jointly reduce robustness under spatial out-of-distribution (OOD) shifts and visual distractors. In this work, we propose SlotFlow, an object-centric flow matching policy for robust visuomotor manipulation. SlotFlow decouples scene observations into semantic ("what") features and lightweight image-plane spatial ("where") cues to provide object-aware policy conditioning and current-state grounding. The semantic representation suppresses irrelevant background correlations, while the spatial cue improves adaptation to shifted object configurations. Extensive simulation and real-world experiments demonstrate improved robustness under visual distractors and severe spatial perturbations while preserving the low-step inference efficiency of A2A. Controlled initialization and perception ablations further identify object-centric grounding as a major source of the gains and show that it complements, rather than replaces, useful historical motion priors.

著者のコメント

Accepted to the 10th Conference on Robot Learning (CoRL 2026)

arXiv ID: 2609.24155 / 要約の誤りについて