arXiv論文メモ
新着一覧
cs.RO · 査読状況未確認

絶えず変わる状況で監視・記憶・行動する双腕ロボット

Watch, Recall, Act: Always-On Robots in Concurrent Embodied Streams

Ding Yi, Peiwen Sun, Chenchu Rong, Jianan Wang, Xili Dai, Xiangyu Yue, Xi Lin

この論文をやさしく読む

ひとことで言うと

指示や周囲の状況が変わり続ける中で、過去の行動を参照しながら両腕を同時に動かすロボットの方法を提案した。

何に役立つ?

考えられる用途は、途中で指示が変わる双腕作業の設計や評価である。要旨で実証されたのは、作成した複合タスクでの性能である。

この研究の面白いところ

監視・記憶・身体状態の情報を非同期に更新し、行動を止めずに基盤モデルへ渡す構成と、遠隔操作データの作成過程をラベル付けに使う点が特徴である。

どこまで分かった?

複合タスクでは45%で、主要比較対象の最高値28%を上回った。要旨には、別の実環境や長期運用での性能は示されていない。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

常時稼働するロボットが向き合うのは、リセットされない連続した情報の流れである。指示は届いたり失効したりし、場面は変化し、自らの過去の行動も次に考えるべき状況を変える。現在の行動モデルは、固定された指示、作業中の介入なし、1段階の推論という逆の前提で設計されている。終わりのない環境では、ロボットはライブの流れを監視してずっと先で必要になる手掛かりを捉え、遠い過去の自分の行動を思い出し、両腕を並行して動かしながら行動しなければならない。 本研究は、意図的に単純に設計したストリーミング方策ARMS(Always-on Robot in Multi-modal Streams)を提示する。事前学習済みの単一のπ0.5基盤モデルに3つの軽量モジュールを加え、ライブの知覚情報、身体状態、自身の過去の行動を、行動前に基盤モデルが参照する文脈に変える。各モジュールは文脈を非同期で更新するため、監視や想起が行動を止めず、両腕を同時に動かせる。新しい機構を考案するのではなく、学習した文脈供給モジュールと、どちらの腕がいつ何をしたかを記録する行動に基づく自己履歴を統合する。 追加の注釈なしで学習させるため、実際の双腕遠隔操作から、段階的な構築スクリプトそのものが各モジュールのラベルを付けるARMS Datasetを作成した。このデータで学習したARMSは複合タスクで45%を達成し、主要な比較対象4種のうち最も良いものの28%を上回った。構成要素を除く実験では、記憶モジュール、身体状態の出力部、非同期の並行処理がそれぞれ必要であることを確認した。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-23(UTC)
最新改訂
2026-09-23 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

An always-on robot faces an endless stream that never resets: instructions arrive and lapse, the scene changes, and its own past actions reshape what it must reason about. Today's action models are built for the opposite: a fixed instruction, no mid-task intervention, single-step reasoning. In an open-ended world a robot must watch a live stream for far-future cues, recall its own far-past actions, and act on them under dual-arm concurrency. We present ARMS (Always-on Robot in Multi-modal Streams), a deliberately simple streaming policy: a single pretrained $\pi$0.5 backbone augmented by three lightweight modules that turn live perception, embodied states, and the robot's own past actions into context the backbone reads before it acts. The modules update this context asynchronously, so watching and recalling never block acting and the two arms act at once. Rather than inventing new mechanisms, ARMS integrates these learned context providers with an agent-causal self-history that logs which arm did what, and when. To supervise them without extra annotation, we build ARMS Dataset, whose staged construction script itself labels every module from real dual-arm teleoperation. Trained on it, ARMS reaches 45% on the combined task against 28% for the strongest of our four main baselines, and ablations confirm the memory module, the embodied-state head, and asynchronous concurrency are each necessary.

著者のコメント

11 pages, 3 figures. Accepted to the 10th Conference on Robot Learning (CoRL 2026)

arXiv ID: 2609.28429 / 要約の誤りについて