arXiv論文メモ
新着一覧
cs.RO / cs.AI · 査読状況未確認

水中の重量物回収を予測する複数視点の世界モデル

Underwater C3-JEPA: An Object-Centric Cross-View World Model for ROV Salvage

Yuncong Yang, Jinlong Li, Yulong Xue, Feng Wu, Chunwen Zhang, Lei Qiao, Xuyang Wang

この論文をやさしく読む

ひとことで言うと

複数の水中カメラ映像と操縦信号から、ROVが物体に触れた後の状態を予測するモデル。

何に役立つ?

水中ROVの重量物回収で、候補操作の評価や行動モデルの学習に使える可能性がある。要旨は表現評価と実際の水中映像での検証を報告する。

この研究の面白いところ

接触センサーなしで、カメラ間の情報を統合する。実映像で、伏せたカメラからの物体状態の再現が直前状態を使う基準より良かった。

どこまで分かった?

実際の水中映像で予測を検証したが、実機が自律的に回収作業を成功させたという結果は要旨にない。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

近距離で重量物を回収する水中遠隔操作機(ROV)向けに、物体を中心に据えた複数視点の予測世界モデルUnderwater C³-JEPAを提示する。名称は、視点をまたぐこと、制御を条件にすること、文脈を拡張することを表す。接触センサーを用いず、同期した複数視点のRGB画像と機体の制御信号から、接触による相互作用や機体の流体力学的な遅れの下で、作業対象の状態が潜在空間でどう変化するかを予測する。C³-JEPAは複数のカメラ観測を作業対象と文脈のトークンに符号化し、学習から除いた視点への注意機構でカメラ間の情報を統合し、制御を条件に将来の状態を直接予測する。弱い対応付けによって、少ない注釈費用で対象物と把持器を定め、SIGRegで幾何学的な表現を鮮明にする。実験では、学習した表現が、画像の再構成を用いない潜在表現の比較手法よりも、後段の評価器へ作業に関係する情報を大幅に多く伝える一方、予測器は軽量に保たれた。得られた予測インターフェースは、モデル予測制御(MPC)の候補評価と、想像した展開に基づく行動エージェントの学習を支援する。実際の水中映像での検証では、同じ構成が伏せたカメラから見た物体状態を再現し、直前状態をそのまま使う基準を上回ったため、この方法がシミュレーション外にも移ることが示された。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-24(UTC)
最新改訂
2026-09-24 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

We present Underwater C$^{3}$-JEPA (cross-view, control-conditioned, context-extended), an object-centric multi-view predictive world model for near-field heavy-load underwater ROV salvage. Without contact sensors, it predicts in latent space how the task-object state evolves through contact interaction and under the hydrodynamic lag of the vehicle, from synchronized multi-view RGB observations and vehicle control signals. C$^{3}$-JEPA encodes multi-camera observations into task-object and context tokens, fuses cross-camera evidence through held-out-view attention, and directly predicts future states conditioned on control. Weak binding anchors the target and gripper at low annotation cost, while SIGReg sharpens the geometric representation. Experiments show that the learned representation transfers substantially more task-relevant information to downstream probes than a reconstruction-free latent baseline, while keeping the predictor lightweight. The resulting predictive interface supports model-predictive-control (MPC) candidate evaluation and imagined-rollout behavior-agent training. Validation on real underwater video shows the same architecture recovering a withheld camera's object state and staying ahead of persistence, so the recipe transfers beyond simulation.

著者のコメント

Submitted to the IEEE for possible publication. 12 pages, 14 figures

arXiv ID: 2609.30214 / 要約の誤りについて