arXiv論文メモ
新着一覧
cs.RO · 査読状況未確認

双腕移動ロボットの動作を時間分割して自動で説明する

Automatic Labelling for Bimanual Mobile Manipulation

Yupu Lu, Jia Pan

この論文をやさしく読む

ひとことで言うと

双腕ロボットの記録を動作の区切りに分け、それぞれで台車と左右の腕が何をしているかを自動で説明する仕組みです。

何に役立つ?

考えられる用途は、長いロボット作業の学習データへの注釈付けや、途中の状態の確認です。要旨は注釈の評価を報告しており、学習後の作業成功率の改善は示していません。

この研究の面白いところ

動作の時間境界を軌道から決めてから、視覚言語モデルに意味を説明させます。左右の腕を無理に同じタイミングにそろえずに扱います。

どこまで分かった?

87.4%は選択したタスクでの出力の反復一致で、正解率ではありません。参加者評価も受容の割合であり、腕のラベルの受容は本体や時間分割より低い結果でした。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

意味のあるサブタスクのラベルは、長期的な行動方策に有用な文脈を与えられるが、信頼できる時間的境界と幅広い意味的説明を自動で特定して注釈を付けることは依然として難しい。本研究では、時間的位置の特定を決定論的な軌道解析に、意味の解釈を視覚言語(VL)推論に割り当てる自動ラベル付けパイプラインを提示する。このパイプラインは、同期した運動学的信号を複数の局面に分割し、各局面に限定したVL推論で内容を説明し、台車、左腕、右腕の動作について出力を集約する。 主にGalaxeaの実機による双腕移動操作29タスクで評価した。まず、選択したタスクでVL推論を3回繰り返したところ、87.4%で同じ出力値が得られた。続いて、9人の参加者が29タスクすべてについてラベル付き局面を確認して判断し、時間分割は90.5%、本体のラベルは90.7%、腕のラベルは78.7%で肯定的に受け入れられた。これらの結果は、分割とVLを組み合わせる設計が、左右の腕が非同期に動く振る舞いを保持しながら構造化された注釈を生成できることを示している。これは、より豊かな意味的サブタスクの同定と、状態に基づく検証の基盤を提供する。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-21(UTC)
最新改訂
2026-09-21 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Semantically meaningful subtask labels can provide useful contexts for long-horizon policies, but automatically identifying both reliable temporal boundaries and broad semantic descriptions for annotations remains difficult. We present an automatic labelling pipeline that assigns temporal localisation to deterministic trajectory analysis and semantic interpretation to vision-language (VL) reasoning. The pipeline segments synchronised kinematic signals into phases, performs phase-localised VL reasoning to describe the contents, and aggregates the outputs for the base, left arm, and right arm actions. We evaluate this pipeline primarily on 29 real Galaxea bimanual mobile-manipulation tasks. Repeating the VL reasoning three times first produces the same output value for 87.4% on selected tasks. A review by nine participants across all 29 tasks then judgements on the labelled phases and shows positive acceptance of temporal divisions (90.5%), body labels (90.7%), and arm labels (78.7%). The results indicate that the segmentation-VL design can produce structured annotations while preserving asynchronous bimanual behaviour, providing a basis for richer semantic subtask identification and state-based verification.

arXiv ID: 2609.24059 / 要約の誤りについて