arXiv論文メモ
新着一覧
cs.CV / cs.RO · 査読状況未確認

ロボット動作を方向と大きさに分けてトークン化

Direction-Scale Decomposition in Action Representation: Rethinking What to Tokenize for Vision-Language-Action Models

Yufei Duan, Hang Yin, Alberta Longhini, Chao Tang, Danica Kragic

この論文をやさしく読む

ひとことで言うと

ロボットの動作を方向と大きさに分けて表現し、視覚・言語・動作モデルの操作成功率を調べた研究です。

何に役立つ?

異なる速度やデータセットのロボット操作例を混ぜて学習する際、動作トークンの作り方を検討できます。

この研究の面白いところ

混合データでのSimplerEnv評価では、従来のBINより全体成功率が10.3ポイント高く、実機でも改善が見られました。

どこまで分かった?

評価はLIBERO、SimplerEnv、指定した実機操作での結果です。多様なデータを混ぜたときの性能低下の緩和は可能性として述べています。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

離散トークンを用いる視覚・言語・動作(VLA)学習では、動作の表現が中心的な役割を果たすが、十分に検討されていない。従来の姿勢増分による表現では、動作トークンが実行速度やデータセット固有の正規化に敏感で、実演やデータセットをまたいで共通する幾何学的構造が見えにくくなり得る。本研究は、トークン化の前に並進と回転の増分を方向成分と大きさ成分へ分ける動作表現、Direction-Scale Decomposition(DSD)を導入する。DSDは動きの方向を切り出し、大きさは別のスケールチャンネルに保持する。 均等なビン分割(BIN)と、Bスプラインに基づくトークン化法BEASTを用い、単一データセットと混合データセットによる学習の両方について、シミュレーションと実機の物体操作でDSDを評価した。LIBEROでは、どちらのトークン化法でも平均成功率が上がった。SimplerEnvでは、混合データセットで学習したDSD-BINの全体成功率がBINを10.3パーセントポイント上回った。実機ロボットの実験でも、ロボット分野での事前学習の有無を問わず改善が見られた。これらの結果は、DSDが離散トークン型VLAモデルの有効な動作表現であり、大規模で多様なデータセットの混合学習で起こる性能低下を和らげる可能性を示す。追加資料は著者らのプロジェクトページにある。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-24(UTC)
最新改訂
2026-09-24 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Action representation plays a central role in discrete-token vision-language-action (VLA) learning but remains underexamined. Under conventional pose-increment representations, action tokens are sensitive to execution speed and dataset-specific normalization, potentially obscuring geometric structure shared across demonstrations and datasets. We introduce Direction-Scale Decomposition (DSD), an action representation that decomposes translation and rotation increments into direction and scale components before tokenization. DSD isolates motion direction while retaining magnitudes in separate scale channels. We evaluate DSD with uniform binning (BIN) and BEAST, a B-spline-based tokenizer, in simulation and real-world manipulation under both single-dataset and mixed-dataset training. On LIBERO, DSD improves average success rates with both tokenizers. On SimplerEnv, DSD-BIN outperforms BIN by 10.3 percentage points in overall success rate under mixed-dataset training. Real-robot experiments further show gains both with and without robotics pretraining. These results support DSD as an effective action representation for discrete-token VLA models and suggest its potential to mitigate performance degradation when training on large and diverse dataset mixtures. Our project page with additional resources is available at https://vla-dsd.github.io/

arXiv ID: 2609.28865 / 要約の誤りについて