arXiv論文メモ
新着一覧
cs.AI / cs.LG · 査読状況未確認

Decision Titanでオフライン強化学習の長期記憶を調べる

Decision Titan: Test-Time Training for Long-Term Memory in Offline Reinforcement Learning

Jude Waide, Robert Lieck

この論文をやさしく読む

ひとことで言うと

過去の情報をニューラルネットワークのパラメータへ書き込む仕組みを、オフライン強化学習に組み込み、長い時間差を越えて覚えられるかを調べます。

何に役立つ?

長い履歴が必要な意思決定モデルを設計する際、記憶機構だけでなく時刻や情報の表し方を検討する参考になります。評価は記憶課題向けのX-Maze環境で行われています。

この研究の面白いところ

コンテキスト窓より20倍長い依存関係を学習したことに加え、内部のゲート値を可視化して記憶の働きを調べています。

どこまで分かった?

訓練時の1.7倍の長さへの汎化は報告されていますが、時刻埋め込みに依存します。長期情報の学習も符号化方法に左右されるため、記憶機構を追加すれば常に長期依存を扱えるという結果ではありません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

長期依存関係は、AI分野の逐次意思決定において依然として大きな課題である。RNNは勾配消失とベクトル型の隠れ状態の表現力の限界に悩まされる一方、Transformerに基づくモデルは、注意機構の計算量が2次で増えることに制約される。近年の研究は、訓練時とテスト時の両方で勾配降下法を使い、ニューラルネットワークのパラメータにエピソード記憶を保存するTest-Time Training(TTT)の枠組みで、この問題に対処することを提案している。この手法は自然言語処理の分野で成功を収めている。しかし、著者らの知る限り、強化学習(RL)への適用も、その記憶が実際にどのように機能するかを解析した研究も、まだ存在しない。 本論文では、Decision TransformerにTTT層を追加したDecision Titanを用いて、オフラインRLにおけるTTTの可能性を調べる。系列的な記憶をテストするためにT-Mazeを拡張したX-Maze環境で、モデルの性能と性質を解析し、ゲート値を時間に沿って可視化することで、記憶機構がどう学習するかを調べる。主要な知見は、Decision Titanがコンテキスト窓の20倍の長さに及ぶ長期依存関係を学習でき、訓練データの1.7倍の長さに汎化することである。ただし重要な点として、時間方向の汎化は用いる時刻埋め込みに依存し、長期依存関係を学習する能力は、関係する情報をどのように符号化するかに依存する。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-10-01(UTC)
最新改訂
2026-10-01 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Long-term dependencies remain a major challenge for sequential decision-making in the field of AI: RNNs suffer from vanishing gradients and the limited expressivity of vector-based hidden states, whilst Transformer-based models are limited by the quadratic scaling of attention. Recent work has proposed tackling this problem with the Test-Time Training (TTT) framework, which stores episodic memories in the parameters of a neural network through gradient descent at both train and test-time. This approach has seen success in the domain of Natural Language Processing, however, to the best of our knowledge it has not yet been applied to the domain of Reinforcement Learning (RL), nor has there been a study analysing how this memory practically functions. In this paper, we study the potential of the TTT framework for offline RL by augmenting a Decision Transformer with TTT layers, dubbed the Decision Titan. We analyse performance and properties of the model in the X-Maze environment, an extension of T-Maze designed to test sequential memory, and investigate how the memory mechanism learns by visualising gate values over time. Our key findings are that Decision Titan can learn long-term dependencies with ranges 20x longer than the context window, generalises to lengths 1.7x the training data, but crucially temporal generalisation depends on the time embeddings used, and the ability to learn long-term dependencies depends on how the relevant information is encoded.

著者のコメント

Accepted at ICML 2026 Workshop on Decision-Making from Offline Datasets to Online Adaptation: Black-Box Optimization to Reinforcement Learning

arXiv ID: 2610.01513 / 要約の誤りについて