arXiv論文メモ
新着一覧
cs.RO · 査読状況未確認

見えない目標を探索して運ぶロボットの能動的な操作手法

From Passive Execution to Active Exploration: Agentic Embodied Manipulation in Realistic Environments

Shilin Ma, Chubin Zhang, Xulong Bai, Zifeng Gao, Shiyi Zhang, Yansong Tang

この論文をやさしく読む

ひとことで言うと

最初は見えない対象物をロボットが自ら探し、見つけてから指定場所へ置くための操作方法です。

何に役立つ?

手掛かりを読み取り、視覚情報を更新しながら物を探して扱うロボットの設計に役立つと考えられます。評価は要旨にある探して置く課題です。

この研究の面白いところ

計画、知覚、実行の三つの部分を協調させ、動作中にも視覚情報を細かく取り直します。

どこまで分かった?

要旨には成功率や比較対象などの定量的な結果が記載されていません。すべての実環境に一般化したとは述べられていません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

近年のエージェントシステムの進歩は、身体を持つロボットによる長い手順の操作能力を大きく高めた。しかし、多くの既存の枠組みは受動的に指示を実行する方式にとどまり、文章による意味的な手掛かり、気を散らす物体、最初は見えない目標を含む現実の場面への適用を制限している。この差を埋めるため、あらかじめ決まった指示を実行するだけでなく、ロボットが環境と動的に相互作用する、エージェントに基づく能動的探索の枠組みを提案する。具体的には、高水準の課題を推論する計画モジュール、視覚的な場面を理解する知覚モジュール、低水準の操作を担う実行モジュールという三つの協調する部分からなる。この設計により、ロボットは課題に関係する情報を能動的に取得し、環境からの反応に応じて行動を変え、一部しか観測できない状況でも操作課題を完了できる。さらに、視覚フィードバックと技能の実行を緊密に組み合わせる、細かな知覚・実行の交互処理を導入し、探索の頑健性を高める。現実的な『探して置く』課題で手法を評価し、操作前に目標物を能動的に発見する必要がある難しい環境での有効性を示した。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-24(UTC)
最新改訂
2026-09-24 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Recent advances in agentic systems have substantially enhanced the long-horizon capability of embodied manipulation. However, many existing frameworks still follow a passive execution paradigm, which limits their applicability to real-world scenarios involving textual semantic cues, distractors, and initially invisible targets. To bridge this gap, we propose an agent-based active exploration framework that enables robots to dynamically interact with the environment rather than merely execute predefined instructions. Specifically, our framework consists of three collaborative modules: a planning module for high-level task reasoning, a perception module for visual scene understanding, and an execution module for low-level manipulation. This design allows the robot to actively acquire task-relevant information, adapt its behavior based on environmental feedback, and complete manipulation tasks under partial observability. Furthermore, we introduce a fine-grained perception-execution interleaving strategy, which tightly couples visual feedback with skill execution to improve exploration robustness. We evaluate our method on a realistic Find-and-Place task, demonstrating its effectiveness in challenging environments where target objects must be actively discovered before manipulation.

著者のコメント

1st Place in the CVPR 2026 GigaBrain Challenge

arXiv ID: 2609.29091 / 要約の誤りについて