検索されない記憶の有用性を測るために検索へ介入する
Causal Memory Policy: Making Memory Utility Identifiable by Intervening on Retrieval
この論文をやさしく読む
ひとことで言うと
一度も取り出されない記憶は役に立つか評価できないため、既知の確率で記憶を検索に混ぜて、その効果を測る方法です。
何に役立つ?
LLMの記憶管理で、使われない記憶を単に不要と判定してしまう問題を調べるために役立ちます。
この研究の面白いところ
記憶を残す・消す操作だけでなく、検索される機会自体を操作します。一方で、有用性を測れれば将来の保持判断も解決するわけではないことまで示しています。
どこまで分かった?
識別は改善しましたが、未見の問合せでの価値を集約スコアから予測する問題は残っています。要旨では評価した集約方法の範囲を超える一般的不可能性までは述べていません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
記憶を拡張した大規模言語モデルは、どの記憶を保持するかを決める必要があり、最近のシステムは各記憶が課題性能へ与える効果を推定して判断する。しかし、これらの推定は検索された記憶に全面的に依存する。一度も検索されない記憶については、保存側への介入を行っても同じ結果となり、その有用性は識別できない。これは検索段階の正値性の違反であり、記憶操作だけを見る診断では見えない。 検索そのものに介入して識別を回復する因果的枠組み、Causal Memory Policy(CMP)を導入する。既知の選択確率で標本化した記憶のために、コンテキストの固定数の枠を確保する。CMPは、均衡化された割付設計のもとで、自己正規化逆確率重み付けを使って記憶の有用性を推定する。記憶の有用性が検索を介して因果的に分解されること、推定量の不偏性と厳密な分散、不可逆操作のもとでの最適な意思決定規則を証明する。 実験では、LongMemEvalで必要な記憶の54%、LoCoMoで67%について識別が失敗し、この失敗は運用中の記憶システムでも続く。CMPは必要な記憶と不要な記憶の識別AUCを0.54から0.66へ改善する。最後に、有用性を識別するだけでは保持の判断に不十分であることを示す。問合せごとの有用性は、それを推定した問合せではAUC 0.78に達するが、保持方策が利用できるどの集約方法も、未見の問合せにおける記憶の価値を予測できない。コードは https://anonymous.4open.science/r/cmp-release-D0C3/ で公開している。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-10-01(UTC)
- 最新改訂
- 2026-10-01 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-10-01 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Memory-augmented large language models must decide which memories to retain, and recent systems do so by estimating each memory's effect on task performance. However, these estimates rely entirely on retrieved memories. When a memory is never retrieved, store-level interventions produce identical outcomes, leaving its utility unidentified. This is a retrieval-level positivity violation, invisible to diagnostics that examine only memory operations. We introduce Causal Memory Policy (CMP), a causal framework that restores identification by intervening on retrieval itself, reserving a fixed number of context slots for memories sampled with known propensities. CMP estimates memory utility by self-normalized inverse propensity weighting under a balanced assignment design. We prove the causal factorization of memory utility through retrieval, the unbiasedness and exact variance of the estimator, and the optimal decision rule under irreversible operations. Empirically, identification fails for 54% of required memories on LongMemEval and 67% on LoCoMo, and the failure persists in a deployed memory system. CMP improves discrimination between required and non-required memories from 0.54 to 0.66 AUC. Finally, we show that identified memory utility alone is insufficient for retention decisions: per-query utility reaches 0.78 AUC on the query for which it is estimated, yet no aggregation available to a retention policy predicts a memory's value on unseen queries. Code is available at: https://anonymous.4open.science/r/cmp-release-D0C3/.
arXiv ID: 2610.02070 / 要約の誤りについて