arXiv論文メモ
新着一覧
cs.RO · 査読状況未確認

似た動作を近いトークンにするロボット行動の符号化

Behavior-Aligned Action Tokenization for Robot Policy Learning

Junbo Dong, Ze Chen, Zhendong Xie, Junjie Li, Lixin Xu, Xuemin Chi, Yiming Song, Zhaoyuan Ma

この論文をやさしく読む

ひとことで言うと

タイミングが違っても似たロボット動作を近い記号で表し、操作方策を学びやすくする手法。

何に役立つ?

考えられる用途は、複数の課題で共通する動きの情報を再利用するロボット操作学習である。要旨ではシミュレーションと二つの実ロボット課題で評価した。

この研究の面白いところ

動作の対応を揃えると方策の成功率は上がったが、記録した軌跡の再生成功率は下がったという差も報告している。

どこまで分かった?

成功率の数値は選ばれたシミュレーション課題の平均であり、実ロボット課題での具体的な数値は要旨にない。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

自己回帰型のロボット方策は、観測から離散的な行動トークンを予測することで連続的な制御を学ぶ。異なる課題でも局所的な動きは共通することが多いが、既存のトークン化手法では実演間の行動の対応関係を明示的に教えることが少ない。そのため、似た動きでもタイミングが異なると共通の表現を持てないことがある。本研究は、Soft-DTWを使って対応する行動のまとまりを選び、再構成と同時に量子化後の座標を揃えるBehavior-Aligned Action Tokenization(BAAT)を提案する。この目的関数は、実行可能な行動の詳細を残しつつ、課題をまたぐ似た動きを近い量子化表現へ配置するよう促す。履歴を条件とする拡散デコーダーがトークンから連続的な行動のまとまりを再構成し、そのトークンの予測を下流の自己回帰方策が学ぶ。 三つのシミュレーションベンチマークから選んだ課題と、二つの実ロボット課題で評価した。BAATのシミュレーションでの平均成功率は約45.2%で、OATを約7.2パーセントポイント上回った。条件を揃えたLIBERO-Allでの対応付けの要素除去実験では、方策の成功率は70.2%から79.0%へ上がる一方、軌跡再生の成功率は下がった。これらの結果は、行動の対応関係を教師信号として共通する動作構造を行動トークン内に整理することが、下流のロボット方策の学習改善につながることを支持する。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-23(UTC)
最新改訂
2026-09-23 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Autoregressive robot policies learn continuous control by predicting discrete action tokens from observations. Different tasks often share local motions, yet behavioral correspondence across demonstrations receives limited explicit supervision in existing tokenizers. Motions with different timing can therefore lack a shared representation despite following similar patterns. We propose Behavior-Aligned Action Tokenization (BAAT), which uses soft dynamic time warping (Soft-DTW) to select corresponding action chunks and aligns their quantized coordinates jointly with reconstruction. This objective encourages similar motions across tasks to occupy nearby quantized representations while retaining executable action detail. A history-conditioned diffusion decoder reconstructs continuous action chunks from these tokens, and a downstream autoregressive policy learns to predict them. We evaluate BAAT on selected tasks from three simulation benchmarks and two real robot tasks. BAAT achieves a mean simulation success rate of approximately 45.2%, exceeding OAT by approximately 7.2 percentage points. In the controlled LIBERO-All alignment ablation, policy success rises from 70.2% to 79.0% while trajectory replay success decreases. These results support behavioral correspondence as supervision for organizing shared motion structure in action tokenizers and improving downstream robot policy learning.

arXiv ID: 2609.27513 / 要約の誤りについて