arXiv論文メモ
新着一覧
cs.RO / cs.LG · 査読状況未確認

行動トークンの再構成誤差とロボット操作性能を比較

Beyond Reconstruction Error: Analytical and Data-Driven Action Tokenization for Autoregressive Vision-Language-Action Models

Yuxin Yang, Gaohan He, Changxue Guan, Hangming Liu

この論文をやさしく読む

ひとことで言うと

ロボットの行動を離散トークンで表す方法は、元の動きを再現する誤差だけでは選べず、実際の制御性能も見るべきだと示した研究。

何に役立つ?

自己回帰型VLAモデルの行動表現を選ぶ際、再構成誤差に加えて予測しやすさや閉ループ制御の成功率を評価する指針になる。

この研究の面白いところ

PCAはTemporal-DCTより再構成誤差が小さい一方、三つの乱数種を平均した既知課題の成功率は3.0ポイント低かった。さらに一つの乱数種では順位が逆で、単一指標だけの選択の難しさが見える。

どこまで分かった?

試行はLIBEROで3,500回行われたが、乱数種によって方策順位が変わる点に注意が必要である。オートエンコーダーの比較は乱数種42にそろえた実験として報告され、一般的な実機性能は要旨からは分からない。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

離散的な行動トークンへの変換は、自己回帰型の視覚・言語・行動モデル(VLA)の中心的な要素だが、行動の表現は主に再構成の正確さによって評価されがちである。本研究では、固定された解析的表現、データに基づく線形表現、非線形のニューラル表現を、共通のトークン化方式の下で比較し、閉ループ制御に本当に重要な表現の性質を調べる。レート歪み解析、系列モデル化の診断、LIBEROでの3,500回の試行を通じて、表現の順位は評価基準によって変わった。主成分分析(PCA)は時間方向の離散コサイン変換(Temporal-DCT)より名目上の再構成誤差が小さいが、トークン系列は予測しにくく、方策学習の三つの乱数種を平均した既知課題の成功率は3.0ポイント低かった。ただし、一つの乱数種では方策の順位が逆転した。乱数種42にそろえた比較実験では、オートエンコーダーが再構成誤差をさらに下げたものの、最も強い方策にはならず、離散トークンの変化に対してもより敏感だった。これらの結果は、再構成の正確さだけでは自己回帰型制御の行動表現を確実に選べないことを示し、幾何学的な忠実度、系列の予測しやすさ、デコーダーの安定性、閉ループでの性能を併せて評価する必要性を示している。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-22(UTC)
最新改訂
2026-09-22 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Discrete action tokenization is central to autoregressive vision-language-action (VLA) models, yet action representations are often evaluated primarily through reconstruction fidelity. We ask which representation properties actually matter for closed-loop control by comparing fixed analytical, data-driven linear, and nonlinear neural representations under a unified tokenization interface. Across rate-distortion analysis, sequence-modeling diagnostics, and 3,500 LIBERO rollouts, representation rankings change with the evaluation criterion. PCA achieves lower nominal reconstruction error than Temporal-DCT, but produces less predictable token sequences and 3.0 percentage points lower mean seen-task success across three policy-training seeds, with the policy ordering reversing in one seed. In a matched seed-42 ablation, an autoencoder further reduces reconstruction error yet does not yield the strongest policy and exhibits greater sensitivity to discrete token perturbations. These findings show that reconstruction fidelity alone cannot reliably select action representations for autoregressive control, motivating joint evaluation of geometric fidelity, sequence predictability, decoder stability, and closed-loop performance.

著者のコメント

5 pages, 1 figure, 4 tables

arXiv ID: 2609.25820 / 要約の誤りについて