arXiv論文メモ
新着一覧
cs.RO / cs.CV · 査読状況未確認

ロボットの経験選択を成功率だけで評価しない監査方法

Beyond End-Task Success: How to Audit Visual Experience Retrieval in Robotics

Eshika Pathak, Leela Krishna

この論文をやさしく読む

ひとことで言うと

ロボットの成功率が高くても、場面に合う経験を上手に選べているとは限りません。経験自体の強さと選び方の良さを切り分ける評価方法です。

何に役立つ?

過去の動作を再利用する方式を比較するとき、どの経験を何回選んだか、最良の一つを使い続けるとどれだけ成功するかを併記する指針になります。

この研究の面白いところ

経験と場面の組全体では成功を予測できる視覚距離でも、一つの場面で候補を選び分ける能力は低いという違いを示しています。

どこまで分かった?

評価範囲は2課題、3再利用機構、5視覚埋め込みと指定のライブラリーサイズです。最良の固定経験は結果を知った後で選んだ比較基準で、事前に分かる設定ではありません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

過去の経験を保存するロボットは、新しい場面でどの経験を再利用するかを選ぶ必要がある。多くのシステムは視覚的な類似性で選択し、多くの評価は選ばれた経験の成功率だけを報告する。しかし、その数値から選択の良し悪しは分からない。広い場面へ転用できる一つの経験を繰り返し使うだけで高得点になることも、好んで選ぶ経験の性能が低いため低得点になることもある。ロボットが再学習よりも再利用によって適応する機会が増えるなか、選択規則ではなく経験ライブラリーの性質を表す得点は、研究分野の次の開発を誤った方向に導く。 本研究では監査方法として、2つの物体操作課題、3つの再利用機構、経験数Kが3、10、50のライブラリーについて、保存されたすべての経験をすべての問い合わせ場面で実行する。選択肢ごとの結果がすべて分かるため、得点が場面別の選択によるのか、ライブラリーの品質によるのかを追跡できる。監査対象の規則は、生の画素からCLIPまで5種類の視覚埋め込みで最近傍距離により選択する。 (1)結果を見て選んだ一つの固定経験だけで、無作為選択と最適な選択を知るオラクルの差の30〜58%を埋められる。場面別の選択が競う余地は、成功率で残り0.07〜0.15である。(2)Kが10以上では、視覚的規則はオラクルの1.5〜3倍、一つの経験に選択を集中させ、得点もその経験の品質に左右される。(3)各経験の選択頻度を保ったまま場面との対応を無作為に入れ替えた場合と有意差があるときには、学習したどの画像ベース方策でも、元の規則のほうが悪い。(4)視覚距離は、ある経験と場面の組が成功するかを高精度に予測できる(AUROCは最大0.96)。しかしKが50では、5種類中4種類の埋め込みで、一つの場面内の候補を順位付けする能力は偶然水準と変わらない(AUROC 0.45〜0.52)。 全組合せの実行は通常困難なので、この監査を、どの研究でも提示できる2つの低費用な報告にまとめる。すなわち、選択された経験の分布と、結果を見て選んだ最良の単一経験の成功率である。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-22(UTC)
最新改訂
2026-09-22 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Robots that store past experiences must select which one to reuse in a new scene. Most systems select by visual similarity, and most evaluations report only the success of the selected experience. That number does not show whether the selection was good: a rule can score well by repeatedly using one broadly transferable experience, or poorly because its preferred experience is weak. Since robots increasingly adapt by reuse rather than retraining, a score that describes the library rather than the rule misleads what the field builds next. We contribute an audit methodology: execute every stored experience in every query scene, over two manipulation tasks, three reuse mechanisms, and libraries of $K=3$, $10$, and $50$. Because every alternative's outcome is known, a score can be traced to per-scene selection or to library quality. The audited rules select by nearest-neighbor distance in five visual embeddings, from raw pixels to CLIP. (1) One fixed experience, chosen with hindsight, captures 30-58% of the gap between random selection and an oracle; per-scene selection competes for the remaining 0.07-0.15 in success rate. (2) At $K\ge10$, visual rules concentrate on one experience 1.5-3 times more than the oracle does, and their scores then follow that experience's quality. (3) Wherever a rule differs significantly from a shuffle that keeps its selection rates but pairs them with scenes at random, the rule is worse, for every learned image policy. (4) Visual distance predicts well whether a given pair will succeed (AUROC up to 0.96), yet ranks the candidates within one scene no better than chance for four of five embeddings at $K=50$ (AUROC 0.45-0.52). Exhaustive execution is usually infeasible, so the audit reduces to two cheap reports any study can give: the distribution of selected experiences, and the success of the best single experience in hindsight.

著者のコメント

Accepted to the IROS 2026 Workshop on Embodied Neuro-Symbolic AI for Reliable and Safe Robotics (ReS AI)

arXiv ID: 2609.26567 / 要約の誤りについて