arXiv論文メモ
新着一覧
cs.RO / cs.LG · 査読状況未確認

ロボットの行動指令の表し方が世界モデルの予測を変える

Robot World Models Are Not Invariant to How the Actions Are Written

Ahmed Karim, Leon Chlon

この論文をやさしく読む

ひとことで言うと

同じロボット指令でも、絶対値と差分という書き方を変えると世界モデルの予測が大きく変わる。

何に役立つ?

ロボットの行動予測モデルを異なる指令形式で使う際の評価や学習法の設計に役立つ。

この研究の面白いところ

二つの表現にはほぼ同じ情報があるのに性能が崩れ、目的関数の平均化と不一致ペナルティで改善する。

どこまで分かった?

評価は記載された三つのデータセットと二つの機体構造などで行われ、PushTでは平均化だけで全ての問題は解消しない。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

ロボットの方策は、関節の目標位置を絶対値で表すか、現在の状態からの差分で表すかのいずれかで訓練される。この選択はロボット学習で実際に必要な設計判断であり、行動を条件にする世界モデルも、その表現方法の影響を暗黙に受ける。本研究は、この影響が深刻であることを示す。一方の表現で訓練した潜在ダイナミクスモデルに、同一の指令軌道をもう一方の表現で与えると性能が崩れる。三つのロボットデータセットと二つの機体構造で検索性能は2.6~13.4倍悪化し、目標を条件とする行動選択は53%から15%に低下する。PushTでは同じ未来についての二つの内部表現がほぼ直交し、コサイン類似度は0.067、最悪の場合は−0.377となる。予測器は徐々に劣化するのではなく、別の問いに答えるようになる。 これは通常の意味での分布変化だけでは説明できない。二つの符号化は関節の入力を合わせれば決定係数R²=0.996で相互に再構成でき、情報は失われていない。研究では、妥当な再パラメータ化を、情報を失う要約やセンサーの交換と区別する検査も提示し、提案した四つの軸のうち三つをこの検査で棄却した。欠陥は行動の入力経路にある。視覚モデルの不変性に関する従来の研究は、切り抜き、画像の揺らぎ、カメラ位置などを扱う一方、指令のパラメータ化は十分に検査していない。 修正策は二つの符号化にわたる平均化だが、平均を取る場所が重要である。目的関数を平均すると、それだけでタスク性能は回復する。出力の平均は確率予測では凹性によって安全だが、方向を値として予測する場合は使えず、正規化した平均が、元の各予測より低い評価になることがある。目的関数の平均化後にも残る問題は、最悪の場合の一致度である。これは0.78のままだが、不一致へのペナルティを加えると0.995に上がる。潜在状態を連続的に予測した場合、最悪の一致度が崩れるか保たれるかの違いになる。PushTでは、平均化だけではその軸は修復されない。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-19(UTC)
最新改訂
2026-09-19 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

A robot policy is trained with one of two action parameterizations: absolute joint targets, or deltas relative to the current state. The choice is a live engineering decision in robot learning, and a world model conditioned on actions inherits it silently. We show the inheritance is catastrophic. A latent dynamics model trained on one parameterization and handed the identical commanded trajectory written in the other collapses: retrieval degrades by 2.6-13.4x across three robot datasets and two morphologies, goal-conditioned action selection falls from 53% to 15%, and on PushT the two beliefs about the same future are near-orthogonal (cos = 0.067, worst case -0.377), so the predictor does not degrade gracefully, it answers a different question. This is not a distribution-shift artifact in the usual sense: the two encodings are mutually reconstructible at R^2 = 0.996 given the joint input, so no information is lost, and we give the test that separates a valid re-parameterization from a lossy summary or a sensor swap. The test rejected three of the four axes we proposed. The defect lives in the action channel, which the invariance literature for visual models does not examine: work there concerns crops, jitter and camera pose, while the parameterization of the commands goes unaudited. The repair is averaging over the two encodings, and where it goes matters. Averaging the objective restores task performance by itself; averaging the outputs, safe for probabilities by concavity, is not available for direction-valued prediction, where the normalized mean can score below every member of the orbit. What objective-averaging leaves behind is the tail: worst-case agreement stays at 0.78, a disagreement penalty closes it to 0.995, and over a latent rollout it is the difference between a worst case that erodes and one that holds. On PushT, averaging alone does not repair the axis.

arXiv ID: 2609.23252 / 要約の誤りについて