GUI操作の経験を学習重みと参照情報へ振り分ける
Not All Experience Belongs in the Weights: Component Routing for Self-Improving GUI Agents
この論文をやさしく読む
ひとことで言うと
画面操作エージェントの経験をひとまとめにせず、内容に応じてモデルへ学習させるか、その場で参照させるかを選ぶ研究です。
何に役立つ?
過去の操作経験をエージェントへ戻す方法を設計する際の判断材料になります。評価した条件では、経験全体を一方の保存先へ送る方法より良い成績を得ています。
この研究の面白いところ
経験の種類に加え、繰り返し現れるか、状態に依存するかという学習前の性質で保存先を選びます。規則の作成に使わなかったモデル系統でも24条件すべてで行き先が一致しました。
どこまで分かった?
報告された比較は3系統の基盤モデルと2環境、3シードの範囲です。平均3.5ポイントの改善に対応する指標の詳しい定義は要旨にありません。コードとデータは公開済みではなく公開予定とされています。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
自己改善するGUIエージェントは、自ら生成した操作軌跡を保存し、ファインチューニングまたは検索によるプロンプトへの取り込みを通じて、その経験をエージェントへ戻す。しかし、この2つの保存先を比較した研究の結論は一致していない。私たちはその原因が経験の単位にあると考える。1本の軌跡には性質の異なる項目がまとめられているため、そのまとまりについての結論は、含まれる項目の構成に依存する。 この問題に対して、第一に、経験を対象の位置特定情報、手順、状態に関する事実、教訓に分割し、各要素をコンテキストまたは重みへ送る「構成要素の振り分け」を導入する。3系統の基盤モデル、2つの環境、3つの乱数シードで、同じ項目について比較した。1つの経験集合に2つの行き先を設けると、位置特定情報と教訓は重みに、手順と状態の事実はコンテキストに置く方が優れていた。 第二に、学習前に測定した2つの性質、反復出現性と状態依存性に基づく規則を当てはめた。この規則は、当てはめに使わなかった基盤モデル系統の24条件すべてで、適切な行き先を再現した。2種類の介入は構成要素を境界に近づけた。この規則による振り分けは、軌跡全体を扱うすべてのベースラインを上回り、各基盤モデルで優れていた単一の保存先と比較しても、平均3.5ポイント高かった。 第三に、学習と、経験の生成側・利用側の違いが、2つの保存先の価値をどう変えるかを明らかにした。同じ構成要素を重みへ書き込むとメモの読み出しが減り、最も頻繁に出現する項目ほど減少が大きかった。コンテキストによる利得は情報の隔たりが大きいほど増え、重みによる利得は方策の隔たりが大きいほど減った。コードとデータは公開予定である。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-10-01(UTC)
- 最新改訂
- 2026-10-01 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-10-01 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Self-improving GUI agents keep the trajectories they produce and return them to the agent, by fine-tuning or by retrieval into the prompt, and studies that compare the two destinations disagree. We attribute this to the unit of experience: a trajectory bundles items with different properties, so a conclusion about the bundle depends on its mix. To address this, (i) we introduce component routing, which splits the experience into locators, procedures, state facts and lessons and sends each component to the context or to the weights, compared on the same items across three backbone families, two environments and three seeds. One pool has two destinations: locators and lessons win in the weights, procedures and state facts in the context. (ii) We fit a rule in two properties measured before any training, recurrence and state-conditionality; it recovers the destination of a held-out backbone family in 24 of 24 cells, two interventions move a component toward the boundary, and routing by the rule beats every whole-trajectory baseline and, by +3.5 points on average, the better single destination of each backbone. (iii) We identify how training and producer-consumer differences change the value of the two destinations: note readout decreases after the same component is written into the weights, most for the items that recur most, context gains increase with the information gap, and weights gains decrease with the policy gap. Code and data will be released.
arXiv ID: 2610.01787 / 要約の誤りについて