arXiv論文メモ
新着一覧
cs.RO / cs.AI · 査読状況未確認

ロボットが自ら視点を変えて情報を集める能力を評価するActiveArena

ActiveArena: Benchmarking and Understanding Active Perception in Robotic Manipulation

Yibo Li, Enshen Zhou, Rui Chen, Yanjun Ding, Mengzhen Liu, Yi Han, Jiabo Zhan, Lipeng Wang, Shanghang Zhang, Lu Sheng

この論文をやさしく読む

ひとことで言うと

ロボットが視点を変え、必要な情報を集めて記憶し、操作に使う能力を測る環境。

何に役立つ?

能動的知覚のモデルを比較し、記憶容量や書き込み方法の効果を調べるのに役立つ。

この研究の面白いところ

35課題と13構成を用意し、受動的な観察だけでは足りない場面で、分布外への汎化も評価する。

どこまで分かった?

要旨に具体的な成功率はない。報告結果は提案したシミュレーターとベンチマークでの比較であり、実機での性能は記載されていない。

v2のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

能動的な知覚と操作は、ロボットが複雑な場面と相互作用するために重要である。既存のベンチマークでは、ロボットが情報を能動的に集め、記憶に保持する能力を十分に評価しにくい。そこで著者らは、視点を制御でき、大規模な作業空間を持つ能動的知覚のシミュレーターActiveArena-Simを基盤として導入する。この上に、細かな五分類にわたる35課題から成り、視覚的な探索と相互作用による情報取得を含むActiveArena-Benchを提案する。各課題は受動的に観察するだけでは解きにくく、複数回の証拠取得と記憶に基づく推論が必要になる。 ベンチマークには豊富な記憶の注釈、標準化した学習データ、場面を分け、未見の妨害物の配置や新しい背景を含む、分布内・分布外の評価手順を備える。さらに能動的知覚での記憶の書き込み、記憶容量、自己受容的な状態、下位課題の教師情報、高水準の計画を統制して調べるため、13種類のモジュール構成からなるActiveArena-VLAを示す。 評価では分布内と分布外の間に大きな性能差が見られた。均一な記憶サンプリング、信頼できる書き込み方策の下での記憶容量の増加、自己受容入力、下位課題の教師情報は、分布外への汎化を改善した。一方、計画器に導かれた記憶管理と意思決定では、少量の記憶だけでも最良の構成に近い性能に達した。ActiveArenaは、能動的な知覚と操作のモデルを開発し、問題点を診断する共通の試験環境を提供する。

v2の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-21(UTC)
最新改訂
2026-09-23 · v2
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Active perception and manipulation are crucial for robots to interact with complex scenes. Existing benchmarks struggle to evaluate how robots effectively acquire and maintain information in memory in an active manner. To this end, we introduce ActiveArena-Sim, an active-perception simulator with controllable viewpoints and large-scale workspaces as the foundation. Built on this, we propose ActiveArena-Bench, which comprises 35 tasks across 5 fine-grained categories, covering visual exploration and interactive information acquisition. Each task is difficult to solve from passive observations alone, requiring multi-round evidence acquisition and memory-based reasoning. The benchmark provides rich memory annotations, standardized training data, and ID/OOD protocols featuring disjoint scenes, unseen distractor configurations, and novel backgrounds. Moreover, we present ActiveArena-VLA, a modular suite of 13 vision-language-action configurations for controlled studies of memory writing, memory capacity, proprioceptive state, subtask supervision, and high-level planning in active perception. Benchmark results reveal a substantial ID-OOD gap: uniform memory sampling, increased memory capacity under reliable write policies, proprioceptive inputs, and subtask supervision improve OOD generalization, while planner-guided memory management and decision-making achieve performance close to the best-performing configuration using only sparse memory. ActiveArena thus provides a unified testbed to develop and diagnose models for active perception and manipulation.

著者のコメント

43 pages. Project page: https://leeibo.github.io/ActiveArena

arXiv ID: 2609.24124 / 要約の誤りについて