次トークン汎関数の推定
Next-token functional estimation
この論文をやさしく読む
ひとことで言うと
時系列の次の観測が新しい種類か、既存データからどれほど離れるか、といった量を推定します。時間的な依存があると通常の1点除外法が失敗する問題に、連続した窓を除く推定法で対処します。
何に役立つ?
依存のある系列で新規性の確率や分類器のテスト誤差を評価するために役立ちます。ここでのトークンは一般の確率変数を指し、言語モデルだけに限定した問題ではありません。
この研究の面白いところ
除外する窓幅が1なら従来のleave-one-outに戻る統一形です。複数の確率過程に対する上界だけでなく、混合Markov連鎖での新規性確率推定に鋭いミニマックス下界も示します。
どこまで分かった?
理論の収束率は定常なβ混合過程でMarton結合を許すなどの仮定の下で成立します。シミュレーションはMarkov連鎖、移動平均、自己回帰過程であり、任意の非定常時系列への保証ではありません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
長さ n+1 の確率変数列の最初の n 点を観測し、観測されていない最後の点と、観測した訓練点の経験測度からなる汎関数を推定したいとする。このような次トークン汎関数には、次のトークンが新規である確率(サプライズ確率)、次のトークンと訓練点との最小距離の裾確率、観測点で学習した分類器のテスト誤差などがある。これらはすべて、従来はleave-one-out法で推定されるが、この方法は時間依存のもとで一致性を持たない。本研究は、各インデックスの後の長さ τ の窓を削除してから経験測度を作るleave-a-window-out推定量を提案する。τ=1ではleave-one-outに戻る。自然な仮定のもとで、Martonカップリングも持つ任意の定常 β-mixing過程に対し、この推定量の誤差がパラメトリック速度で減少することを示す。この結果は、広いクラスの確率過程上のいくつかの自然な汎関数を含む。さらに、混合マルコフ連鎖でサプライズ確率を推定する場合の鋭いミニマックス下界を示し、上界を補完する。マルコフ連鎖、移動平均過程、自己回帰過程のシミュレーションでは、leave-one-out法や加法定数ベースラインが失敗する多くの状況で提案推定量が成功した。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-17(UTC)
- 最新改訂
- 2026-09-17 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-17 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Suppose we observe the first $n$ points of a sequence of random variables having length $n+1$, and wish to estimate a functional of the unobserved final point and the empirical measure of the $n$ observed training points. Such next-token functionals include the probability that the next token is novel (also known as the surprise probability), the tail probability of the minimum distance between the next token and training points, and the test error of a classifier trained on the observed points. All of these quantities are classically estimated by the leave-one-out method, which is inconsistent under temporal dependence. We propose a leave-a-window-out estimator, which deletes a window of length $\tau$ after each index before forming the empirical measure and reduces to leave-one-out at $\tau = 1$. Under natural assumptions, we show that the error of our estimator decays at a parametric rate for any stationary $\beta$-mixing process that also admits a Marton coupling. Our results thus cover several natural functionals on a large class of stochastic processes. We complement these upper bounds with a sharp minimax lower bound for estimating the surprise probability on mixing Markov chains. Simulations on Markov chains, moving-average processes, and autoregressive processes show that our estimator succeeds in many scenarios where leave-one-out and add-constant baselines fail.
arXiv ID: 2609.19529 / 要約の誤りについて