arXiv論文メモ
新着一覧
cs.LG / eess.SP / math.OC · 査読状況未確認

目標を定期更新する確率的反復法の収束解析

A Contraction Framework for Stochastic Operators with Bootstrapping: Application to TD Learning

Ids van der Werf, Sergio Rozada and Antonio G. Marques

この論文をやさしく読む

ひとことで言うと

学習の目標値を何ステップかごとに更新する反復計算について、収束条件と誤差を調べた理論研究です。

何に役立つ?

TD学習などで目標の更新周期やステップ幅を検討する際、収束の見通しを与えると考えられます。要旨には理論的な上界とTD学習のシミュレーションが示されています。

この研究の面白いところ

更新が勾配計算に基づくことを要求せず、標本誤差が反復値とともに増える場合も扱います。既存の複数の結果を一つの収縮性の枠組みから導いています。

どこまで分かった?

幾何学的収束には、固定目標への感度が内部写像の収縮余裕より小さいという条件があります。収束先も固定点そのものとは限らず、その周囲の領域です。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

多くの反復アルゴリズムは、更新対象とは別に固定したコピーを目標として使い、その目標を定期的に更新するブートストラップに依存している。majorize-minimize法、不正確な近接点法、時間差分(TD)学習には、この構造が共通する。しかし、標本に基づく更新と、Kステップごとにしか更新されない目標を組み合わせた場合、既存の収束保証は、線形近似や勾配に基づく内部ステップといった更新の固有の構造や、標本誤差が一様に有界であることに頼っていた。 本研究では、標本に基づく更新をパラメータ空間上の確率的作用素としてモデル化する。これにより、解析を収縮性に関する議論へ還元し、勾配構造を必要とせず、反復値とともに標本誤差が増える場合も許す。この枠組みの下で、独立同分布の標本と任意の目標更新周期Kについて、有限時間の上界を導く。固定目標に対する感度が内部写像の収縮性の余裕より小さければ、反復値は固定点の周囲の領域に、二乗平均平方根の意味で幾何学的に収束することを示す。既存の決定論的な固定目標の収縮結果と確率的勾配型の上界は、この枠組みの特別な場合として導かれる。TD学習のシミュレーションでも、予測した収縮率と、ステップ幅に対する誤差下限の変化が再現された。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-24(UTC)
最新改訂
2026-09-24 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Many iterative algorithms rely on bootstrapping. A variable is updated using a second, frozen copy as a target, which is periodically replaced with the updated variable. Majorize-minimize and inexact proximal-point methods share this structure, as does temporal-difference (TD) learning. However, existing convergence guarantees for scenarios that combine sampled updates with targets refreshed only every $K$ steps rely on the specific structure of the update, such as linear approximation or gradient-based inner steps, and on uniformly bounded sampling error. We instead model the sampled update as a stochastic operator on the parameter space, which reduces the analysis to a contraction argument that needs no gradient structure and allows the sampling error to grow with the iterates. Within this framework, we derive a finite-time bound for i.i.d. samples and any target-update period $K$. We show that the iterates converge geometrically in root mean square to a ball around the fixed point, provided the sensitivity to the frozen target is smaller than the contraction slack of the inner map. Existing deterministic frozen-target contraction and stochastic-gradient-type bounds follow as special cases of our framework, and simulations of TD learning reproduce the predicted contraction rate and scaling of the error floor with the step size.

著者のコメント

5 pages, 1 figure

arXiv ID: 2609.29961 / 要約の誤りについて