arXiv論文メモ
新着一覧
cs.RO · 査読状況未確認

身体を持つエージェントの思い違いを再確認と条件付き巻き戻しで修正

AquaMend: Minimal Re-probing and Conditional Rollback for Latent-Belief Failures in Embodied Agents

Yufan Liu, Shang Luo, Yang Liu, Haoxuan Jia, Feiyu Han, Qian Li, Chen Li, Yingguang Yang, Chongyang Zhang, Hao Zheng, Kefu Xu, Bin Chong

この論文をやさしく読む

ひとことで言うと

ロボットなどが環境について誤った前提を持ったとき、再確認、やり直し、続行を損失に基づいて選びます。

何に役立つ?

感知し直すコストと失敗の損失を比べて、行動の修正方法を決める設計に役立つ可能性があります。要旨の評価は独自シミュレーションです。

この研究の面白いところ

単に巻き戻すのではなく、再確認と続行も同じ枠組みで比較し、最適性が保証される条件を明示しています。

どこまで分かった?

一般の場合に大域最適性の保証はありません。DTTとの損失差はHolm補正後に有意ではなく、後期段階では意思決定時間が増えました。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

物理環境の変化や感知の誤りは、身体を持つエージェントが課題に必要な事柄について抱く信念を無効にし得る。AquaMendは、探索・信念・行動のグラフ上で、再確認、巻き戻し、根拠に基づく続行を比較する。その際、感知、物理的な復旧、修正されない失敗を含む期待損失を目的関数とする。共同の事後分布を使い、条件付きで検出力を確認する一段階の方策を導く。信念ごとに三つの選択肢から最適なものを選べるのは、独立性、分離可能性、完全に問題を解消する確認手段がある場合である。一般の方策には大域的な最適性の保証はない。 独自に構築したシミュレーションのベンチマークで、対になる32場面を評価した結果、AquaMendは32件中28件で復旧し、最初からやり直す方法に比べて平均の総損失を21.6%減らした。意思決定理論に基づくトラブルシューティング(DTT)との対になった損失差は、Holm補正後には統計的に有意ではなかった。全候補を使う要素除去版との比較では、オンラインの意思決定時間が全体で12.3%減ったが、対象外の後期段階では3.4%増えた。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-24(UTC)
最新改訂
2026-09-24 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Physical changes or sensing errors can invalidate embodied agents' task-relevant beliefs. AquaMend compares re-probing, rollback, and supported continuation on a probe-belief-action graph under an expected-loss objective covering sensing, physical recovery, and uncorrected failures. A joint posterior guides a one-step policy with conditional detection-power screening. The per-belief three-way optimum requires independence, separability, and fully resolving probes; the general policy has no global optimality guarantee. Across 32 paired scenarios in a self-constructed simulation benchmark, AquaMend recovers in 28/32 cases and reduces mean complete loss by 21.6% versus restart. Its paired loss difference from decision-theoretic troubleshooting (DTT) is not statistically significant after Holm correction. Against the all-candidate ablation, online decision time decreases by 12.3% overall but increases by 3.4% in the uncovered late stage.

著者のコメント

29 pages, 1 figure. Yufan Liu, Shang Luo, and Yang Liu contributed equally. Corresponding author: Bin Chong

arXiv ID: 2609.28973 / 要約の誤りについて