arXiv論文メモ
新着一覧
cs.LG / cs.AI · 査読状況未確認

環境の事実と作業手順を分けて更新するエージェント

TTSE: A Two-Track Online Self-Evolution Framework

Ruimin Pei, Yongkang Wu, Shangyi Zheng, Yaqing Zhang, Deyang Li, Jianjun Tao, Xinyu Zhang, Xiang Zhang

この論文をやさしく読む

ひとことで言うと

エージェントが覚える内容を、環境について確かめた事実と、タスクを実行する手順に分け、それぞれ更新する方法です。

何に役立つ?

継続して作業するエージェントが、環境を取り違えたり、条件に合わない手順を使ったりする問題を整理するのに役立ちます。長期自律性は研究の目的で、あらゆる環境での達成を実証したわけではありません。

この研究の面白いところ

二種類の知識の分離を、理論上のリスク分解と実際のアブレーション比較の両面から検討しています。既存のスキル更新手法と組み合わせる評価もあります。

どこまで分かった?

優位性は挙げられたベンチマークにおける結果です。SOPBenchとPinchBenchでは3回の独立実行を報告しますが、具体的な改善値や長期間の運用成績は要旨にありません。理論的な優位性にも条件があります。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

大規模言語モデル(LLM)エージェントが継続的に相互作用する環境へ適用されるにつれ、自身の能力を進化させることが長期的な自律性の実現に向けた中心課題となる。現状では、環境に関する知識は、エージェントの継続的な進化の一部ではなく、外から与えられる固定入力として扱われることが多い。強化学習法は通常、環境との相互作用で方策を最適化するが、固定されたタスク分布や単一環境への適応にとどまりがちである。 本論文では、更新される知識をFACTとTIPに分ける、二系統のオンライン自己進化フレームワークTTSE(Two-Track Self-Evolution)を提案する。FACTは相互作用の証拠を通じて信頼性を継続的に検証する環境の事実で、TIPはタスクを条件とする実行手順である。意思決定理論の観点から、エージェントの超過リスクを環境表現のリグレットと条件付き実行のリグレットに分解する。そして、環境を条件とする方策が条件を考慮しない方策を厳密に上回る条件を特徴付け、FACTの識別誤差と条件間の不一致費用によって下流リスクを上から抑える。 実際の評価では、GDPevoでのアブレーション実験により二系統の進化の利点を確認した。代表的なエージェントタスクのベンチマークALFWorldとScienceWorldでも、優れたタスク適応を示す。さらに、TTSEは既存のスキル自己進化手法と広く併用できる。Bayesian-Agentアルゴリズムと組み合わせた際には、単一系統にしたアブレーションとの比較で二系統の利点が確認され、3回の独立した反復でSOPBenchの主要5分野にわたる総合スコアが大きく改善した。最後に、実際のエンドツーエンドタスクのベンチマークPinchBenchでは、検索に基づく情報注入を通じてTTSEを汎用エージェントの枠組みに統合し、3回の独立した実行でベースラインを安定して上回った。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-21(UTC)
最新改訂
2026-09-21 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

As Large Language Model (LLM) agents are applied in continuously interactive environments, driving the evolution of their own capabilities becomes a core problem for achieving long-term autonomy. Currently, environmental knowledge is typically treated as an external fixed input rather than as part of the agent's ongoing evolution. Reinforcement learning methods usually optimize policies through environmental interaction but tend to adapt only to fixed task distributions or single environments. This paper proposes TTSE (Two-Track Self-Evolution), a dual-track online self-evolution framework that separates evolving knowledge into FACT (environmental facts, whose reliability is continuously verified through interaction evidence) and TIP (task-conditioned implementation procedures). From a decision-theoretic perspective, we decompose the agent's excess risk into environment-representation regret and conditional-execution regret, characterize the conditions under which environment-conditioned policies strictly outperform condition-agnostic policies, and bound the downstream risk in terms of FACT identification error and cross-condition mismatch cost. In practice, TTSE's ablation experiments on GDPevo validate the advantage of dual-track evolution. On the classic agent task benchmarks ALFWorld and ScienceWorld, TTSE further demonstrates superior task adaptation. Moreover, TTSE is broadly compatible with existing skill self-evolution methods; combined with the Bayesian-Agent algorithm, a single-track ablation validates the dual-track advantage, substantially improving the aggregate score across the five major domains of SOPBench over three independent repetitions. Finally, on the real end-to-end task benchmark PinchBench, TTSE is integrated into a general agent framework via retrieval-based injection and stably outperforms the baseline across three independent runs.

著者のコメント

20 pages, 2 figures, 18 tables

arXiv ID: 2609.24289 / 要約の誤りについて