arXiv論文メモ
新着一覧
cs.RO · 査読状況未確認

長時間の物体操作を支える四次元の物体記憶

OCC4M: Object-Centric 4D Memory for Spatiotemporal Reasoning in Long-Horizon Manipulation

Jack B. Jedlicki, Tanguy Dieudonné, Heng Yang

この論文をやさしく読む

ひとことで言うと

物体の位置と関係を時間を通じて記憶し、見えなくなった物体を扱うロボット操作を支える手法。

何に役立つ?

容器の入れ替えや視点変更を含む長時間操作で、対象物を選ぶための記憶設計に役立つと考えられる。

この研究の面白いところ

全観測履歴をそのまま視覚言語モデルへ渡す方法と、同じ実行器で比較している。シミュレーション350回と実機20回を報告した。

どこまで分かった?

実機の二段階課題の完了率は45%である。高い成功率の主な数値は指定されたシミュレーション条件と視点移動試験で得られた。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

長時間にわたる物体操作では、見えなくなった物体の位置、時間を通じた同一性、入れ替えられた容器の中身など、現在の視野にない状態を推論する必要がある。OCC4M(「Occam」)は、共通の世界座標系で物体の追跡を継続し、時間、動き、包含関係を明示的に表す、物体中心の四次元記憶である。視覚言語モデルは、この構造化された記憶を照会し、過去の履歴を使わない低水準の実行器に渡す操作対象を選ぶ。 七つのシミュレーション条件、計350エピソードで、OCC4Mの記憶課題成功率は96.6%、一連の操作の成功率は88.9%だった。全観測履歴を用いるGemini 3.7 FlashベースのFrameSampを同じ実行器と組み合わせた場合は、それぞれ54.6%、57.7%だった。視点移動を制御した試験では、視点変更後もOCC4Mは記憶課題で100%、一連の操作で98%の成功率を保ち、全履歴を使うFrameSampはほぼ成功しなかった。固定カメラのFrankaロボットによる20エピソードでは、OCC4Mの記憶と行動の両方の正確率は85%で、FrameSampは文脈長をK=16から全履歴まで変えても最大30%だった。二段階の課題全体の完了率は45%だった。これらの結果は、長時間の操作で持続的な時空間推論を行うための明示的な物体中心記憶を支持する。定性的な動画は https://occ4m-sup.github.io/occ4m-supplementary/ で公開されている。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-23(UTC)
最新改訂
2026-09-23 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Long-horizon manipulation often requires reasoning about state absent from the current view, such as a vanished object's location, temporal identity, or the contents of a shuffled container. We present OCC4M ("Occam"), an object-centric 4D memory that maintains persistent tracks in a shared world frame and explicitly represents temporal, motion, and containment relations. A vision-language model (VLM) queries this structured memory to select actionable targets for history-free low-level execution. Across seven simulation conditions and 350 episodes, OCC4M achieves 96.6% memory success and 88.9% end-to-end success, versus 54.6% and 57.7% for FrameSamp, a raw-history VLM baseline using Gemini 3.7 Flash with the complete observation history and the same executor. In a controlled viewpoint-transfer test, OCC4M maintains 100% memory and 98% end-to-end success after a viewpoint change, while full-history FrameSamp falls to near-zero success. On 20 fixed-camera Franka episodes, OCC4M reaches 85% joint memory accuracy, versus at most 30% for FrameSamp across context sizes from $K=16$ to the complete history, and completes 45% of full two-stage tasks. These results support explicit object-centric memory for persistent spatiotemporal reasoning in long-horizon manipulation. Qualitative videos are available at https://occ4m-sup.github.io/occ4m-supplementary/.

著者のコメント

13 pages, 11 figures, 6 tables. Supplementary videos: https://occ4m-sup.github.io/occ4m-supplementary/

arXiv ID: 2609.28798 / 要約の誤りについて