arXiv論文メモ
新着一覧
cs.AI · 査読状況未確認

エージェントに現在の状況と未完了事項を明示して長期作業を支える

Beyond Memory: Harnessing Long-Horizon Agents with Explicit Belief States

Yu Luo, Jiamin Jiang, Yimin Zuo, Xidao Wen, Rongchen Gao, Yongqian Sun, Shenglin Zhang, Guiyang Liu, Cheng Zhang, Fang Situ, Qi Zhou, Dan Pei

この論文をやさしく読む

ひとことで言うと

エージェントに履歴だけを渡すのではなく、現在の状況の推定と未完了事項を更新し続け、進展が止まった場合には回復処理を行う方法です。

何に役立つ?

長い作業で同じ行動を続けたり、未解決の要件を忘れたりする問題への対処に役立つ可能性があります。

この研究の面白いところ

状況理解の整合性と実際の進捗を別々に確認し、停滞の種類と残る要件に合わせて回復する点が特徴です。

どこまで分かった?

三つの基盤モデルと四つのベンチマークでの比較です。要旨には成功率の具体値や追加計算コストは記載されていません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

大規模言語モデル(LLM)のエージェントは、ますます複雑な課題を担えるようになっているが、やり取りの履歴を記憶へ整理する方法は、現在の世界について一貫した理解を保証しない。本研究では、エージェントの意思決定の文脈として明示的な信念状態を構築し、継続的に維持する、推論時の枠組みPoSを導入する。各信念は、現在の世界の状態の推定と未解決の課題要件を組み合わせ、エージェントがまだ何を知り、何を達成する必要があるかを明示する。 この信念を信頼でき、行動に使える状態に保つため、PoSは整合性を検証し、課題の進捗を監視する。そして、目標への意味のある進展がないまま行動を続ける「信念の罠」を検出する。回復処理は、罠のパターンと未解決の課題要件の種類の両方に合わせて調整する。 実行と診断にまたがる四つのベンチマークでの実験により、PoSは三つのLLM基盤モデルのすべてで、各ベンチマークの総合性能が最も高いことを示す。構成要素を除く実験は、整合性検証と回復処理の重要性を示し、文脈量を増やす実験は、文脈の増大に対する耐性を示す。これらの結果は、履歴の保持や圧縮を超えた長期的な文脈管理の基盤として、信念の構築と継続的な維持を支持する。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-10-01(UTC)
最新改訂
2026-10-01 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Large language model (LLM) agents can now undertake increasingly complex tasks, but the way they organize interaction history into memory does not ensure a coherent understanding of the current world. We introduce PoS, an inference-time framework that constructs and continually maintains explicit belief states as the agent's decision context. Each belief combines an estimate of the current world state with unresolved task requirements, making explicit what the agent still needs to learn and accomplish. To keep this belief reliable and actionable, PoS validates its consistency and monitors task progress to detect Belief Trapping, where the agent continues to act without making meaningful progress toward the goal. Recovery is then tailored to both the trapping pattern and the type of unresolved task requirement. Experiments on four benchmarks spanning execution and diagnosis show that PoS achieves the highest overall performance on every benchmark with all three LLM backbones. Ablations demonstrate the importance of consistency validation and recovery, while context-scaling experiments show resilience to context growth. Together, these results support belief construction and continual maintenance as a foundation for long-horizon context management beyond history retention and compression.

arXiv ID: 2610.01415 / 要約の誤りについて