長い身体化タスクでエージェントの記憶を測るベンチマーク
EmbodiedMemory-Bench: Benchmarking Embodied Memory for Long-Horizon Embodied Tasks
この論文をやさしく読む
ひとことで言うと
環境で長く行動するAIが、過去に見聞きしたことを覚えて後の課題に使えるかを測るベンチマークと記憶システム。
何に役立つ?
身体化エージェントの記憶の弱点を、視覚、状態変化、行動結果、経験からの一般化に分けて評価するのに役立つ。
この研究の面白いところ
四課題群・2,554エピソードを用意し、空間・出来事・場面を整理する外部記憶EMemが、同じ基盤モデルで比較した記憶方式中の総合性能を上回った。
どこまで分かった?
要旨は比較の順位を述べるが、各課題の具体的な成功率は記さない。評価対象以外の長期タスクで同じ効果があるかは分からない。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
長時間にわたる環境との相互作用では、エージェントが観測、行動、変化に遭遇するたびに、環境についての情報を保持し更新し続ける必要がある。しかし、現在のエージェントはこの記憶を安定して維持するのが苦手である。著者らの分析は、細部の視覚記憶の弱さ、変化する世界の状態追跡の不確かさ、行動の結果から分かった世界の状態を記録できないこと、過去の経験からの一般化の弱さ、という四つの問題を指摘する。既存のベンチマークは、長い相互作用の中でこれらの記憶能力を直接評価しない。そこで、四つの課題群にわたる2,554の対話的エピソードから成るEmbodiedMemory-Bench(EMem-Bench)を導入する。エージェントは相互作用の履歴から記憶を作って更新し、後の課題で環境内に行動してその記憶を使う。 さらに、身体化された経験を空間、出来事、場面の記憶に整理する外部記憶システムEmbodied-Memorizer(EMem)と、これらの記憶を管理・利用する80億パラメータの方策EMem-8Bを提示する。多様なオープンソースおよび独自のマルチモーダルLLMと、代表的なマルチモーダル記憶システムを評価した。現行モデルの性能は四つの課題で低く、ばらつきも大きかった。同じ基盤モデルで比べると、EMemは評価した記憶システム中で総合性能が最も高く、オープンソース・独自モデルの双方を改善した。EMem-8Bも元の基盤モデルを上回った。プロジェクトページは要旨に記載されている。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-23(UTC)
- 最新改訂
- 2026-09-23 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-23 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Long-horizon embodied interaction requires agents to retain and continually update information about the environment as they observe, act, and encounter change. Yet current agents struggle to maintain such memory reliably. Our analysis traces this limitation to four key deficiencies: weak fine-grained visual memory, unreliable dynamic world-state tracking, failing to record world state revealed by interaction outcomes, and limited generalization from prior experience. However, existing benchmarks do not directly assess these memory capabilities during long-horizon embodied interaction. To address this gap, we introduce EmbodiedMemory-Bench (EMem-Bench), comprising 2,554 interactive episodes across four task families. EMem-Bench requires agents to build and update memory from interaction history, then use it to complete a later task by acting in the environment. We further present Embodied-Memorizer (EMem), an external memory system that organizes embodied experience into spatial, event, and scene memories. We also train EMem-8B, an 8B policy that manages and uses these memories. We evaluate a diverse range of open-source and proprietary MLLMs and representative multimodal memory systems. Results show that current models remain weak and uneven across the four challenges. Under matched backbones, EMem achieves the best overall performance among the evaluated memory systems and improves both open-source and proprietary models, while EMem-8B further improves over its backbone. Project page: https://zju-omniai.github.io/EmbodiedMemoryBench/
arXiv ID: 2609.28236 / 要約の誤りについて