arXiv論文メモ
新着一覧
cs.RO · 査読状況未確認

ロボットの記憶が行動選択に役立つかを検証する方法

Memory That Changes Action Is Not Memory That Guides It: Counterfactual Auditing of History-Conditioned Robot Policies

Jiajie Zhang, Yankai Xiang, Changhao Chen

この論文をやさしく読む

ひとことで言うと

同じ現在の状態に至る二つの履歴を使い、ロボットの行動が記憶に反応するだけでなく、履歴に合った選択になっているか調べる研究。

何に役立つ?

考えられる用途は、記憶を持つロボット方策の評価。成功率や行動の変化だけでは見落とす、履歴に対する判断の誤りを調べられる。

この研究の面白いところ

共通の現在状態で履歴を入れ替え、各行動を両方の履歴で評価する。実機では記憶により行動が変わっても、完了した9件中5件が誤った目標に達した。

どこまで分かった?

Mem-0課題、内部介入、双腕実機での結果を示す。実機の件数は9件であり、あらゆるロボット方策への一般化までは要旨から判断できない。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

ブロックを元のトレーに戻すロボットは、現在の入力は同じでも、そこに至る二つの履歴によって取るべき行動が異なる状況に遭遇しうる。しかし、記憶を使う方策の評価では、タスクの成功や記憶を変更したときの行動の変化がよく使われる。いずれも、記憶が判断を適切に導いた証拠にはならない。 本研究は、反事実的記憶監査(CMA)という評価手順を提案する。現在の状態が同一であると検証したうえで二つの履歴を組み合わせ、固定した方策に共通の乱数条件で問い合わせ、保存した各行動を両方の履歴の下で評価する。これにより、記憶への感度、履歴に照らした選択の妥当性、対応する状況での物理的価値、履歴の組ごとの信頼性を分けて測れる。 Mem-0のPut Back課題では、監査したすべての組で行動が変わったが、完全に信頼できたのは64組中20組だけだった。後のSwapの判断点では、組になった行動はすべて変わった一方、両方の記憶が同じ分岐を選んだ。システム内部への介入では、履歴バンクを置き換えると行動がその内容に沿って変わり、保護した4096バイトの基準情報を復元すると、注入したバンクの障害で失われたSwap成功率が38.9ポイント回復した。双腕の実機でも記憶は保存される行動を変えたが、完了したPut Back操作9件のうち5件は誤った目標に到達した。ロボットが記憶し反応していても、過去に照らして適切な行動を安定して選べるとは限らないことを示し、CMAをその区別に使う判断単位の監査法として提示する。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-23(UTC)
最新改訂
2026-09-23 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

A robot returning a block to its origin tray may encounter two task-consistent pasts that reconverge to the same current input but warrant different actions. Yet memory-policy evaluations often rely on task success or action change under memory perturbation, neither of which establishes that memory guides the decision. We propose the \textbf{Counterfactual Memory Audit (CMA)}, an evaluation protocol that crosses two histories at a verified-identical present, queries a frozen policy under common randomness, and evaluates each saved action under both pasts. This separates memory sensitivity, warranted choice, matched-world physical value, and per-pair reliability. On Mem-0, every audited Put Back pair changes action, but only $20/64$ pairs are fully reliable; at a later Swap decision, all paired actions change while both memories select the same branch. Native interventions further show closed-loop influence: replacing the history bank redirects behavior toward the replaced content, while restoring a 4096-byte protected anchor recovers $38.9$ points of Swap success lost to injected bank faults. On a dual-arm physical platform, memory changes saved actions, yet five of nine completed Put Back manipulations reach the wrong target. These results show that a robot can remember and react without reliably using memory to choose the behavior its past warrants. CMA provides a decision-level audit for distinguishing these cases.

arXiv ID: 2609.27247 / 要約の誤りについて