arXiv論文メモ
新着一覧
cs.RO / cs.AI · 査読状況未確認

新しいカメラ視点に強いロボット方策のための画像拡張

InfiNoVA: Infinite Novel View Augmentation for Viewpoint Invariant Robot Policies

Sai Puneeth Reddy Gottam, Elmar Rueckert, Vedant Dave

この論文をやさしく読む

ひとことで言うと

複数カメラの実演を3Dで再構成し、未知の視点でも動けるロボット方策を学ぶ。

何に役立つ?

考えられる用途は、実カメラを大量に増やさずに視点変化への頑健性を高めることである。

この研究の面白いところ

状態と行動の対応を保ったまま新視点を描画し、生成画像の作業上重要な誤りを減らす。

どこまで分かった?

実世界の四つの操作課題での比較結果を示した。ほかの作業や環境での成績は要旨に記載がない。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

視覚・言語・行動(VLA)方策は、学習時に見たカメラ視点に強く依存し、未知の視点で使うと性能が大きく落ちることがある。多様な実カメラ視点での実演収集は高価で、それでも視点空間をまばらにしか覆えない。そこで、同期した複数カメラの実演から、幾何学的に整合する高密度な学習用視点を作るデータ拡張法InfiNoVAを提案する。各操作軌跡を時間変化する3D Gaussian表現として再構成し、元の状態と行動の対応を保ちながら、標本抽出したカメラ位置から新しい観測画像を描画する。明示的なシーン表現により、フレームごとの忠実度と時間的な一貫性が向上し、生成型の新視点合成で見られる、作業に重要な部分の幻覚を減らす。実世界の四つの操作課題で、InfiNoVAによって学習した方策は、未知のランダム視点における平均成功率が、VISTAに基づく拡張と拡張なしの方策の双方に比べ5.4倍だった。また、五つの実カメラ視点すべてを直接使って学習する場合よりも成功率が1.7倍高かった。基礎となる方策の構造を変えず、幾何学に基づく高密度な視点拡張でカメラ視点に強いロボット方策を作れることを示す。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-23(UTC)
最新改訂
2026-09-23 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Vision-Language-Action (VLA) policies often rely strongly on the camera viewpoints seen during training, causing substantial performance degradation when deployed from unseen perspectives. Collecting demonstrations from sufficiently diverse physical viewpoints is expensive and still provides only sparse coverage of the viewpoint space. We introduce InfiNoVA, a data-augmentation framework that converts synchronized multi-camera demonstrations into a dense distribution of geometrically consistent training views. InfiNoVA reconstructs each manipulation trajectory as a time-varying 3D Gaussian representation and renders novel observations from sampled camera poses while preserving the original state-action correspondence. This explicit scene representation improves frame-level fidelity and temporal consistency while reducing task-critical hallucinations observed in generative novel-view synthesis. Across four real-world manipulation tasks, policies trained with InfiNoVA achieve 5.4x higher average success under unseen randomized viewpoints than both VISTA-based augmentation and the unaugmented policy. InfiNoVA further achieves 1.7x higher success than training directly on all five physical camera views. These results show that dense, geometrically grounded viewpoint augmentation provides a practical route toward camera-robust robot policies without modifying the underlying policy architecture.

arXiv ID: 2609.27734 / 要約の誤りについて