arXiv論文メモ
新着一覧
cs.LG / cs.AI · 査読状況未確認

現在は正しく答えられる記憶が将来の更新に耐えるか

Correct Now, Insufficient Later: Auditing Update Sufficiency in Context Compression

Guangzhe Zhang

この論文をやさしく読む

ひとことで言うと

今の質問に答えられる圧縮メモでも、後の情報更新に必要な違いを失っている場合があると検証しています。

何に役立つ?

対話履歴やエージェントの記憶を圧縮するとき、現在の正答率だけでは見逃す欠落を点検するために役立ちます。

この研究の面白いところ

現在の答えは同じでも、同一の追加情報を受けた後の答えは異なる履歴を対にします。内容の欠落と出力形式の不備を分け、名前の付け方への依存も別途調べています。

どこまで分かった?

24組の合成履歴を用いた予備評価で、独立した未知データや自然な課題での検証は主張していません。提示区間は信頼区間ではなく有限標本の識別区間です。名前依存の修正も将来の関連性の問題を解決していません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

記憶は、現在の質問には正しく答えられても、将来の更新に必要な区別を捨ててしまうことがある。本研究では、履歴を対にした監査によってこの失敗を調べる。2つの履歴は現在の答えが同じで、共通の将来の更新を受けるが、その後は異なる答えを必要とする。予備評価では、6種類の合成メカニズムにわたる24組の履歴、12の記憶条件、2回の反復、2つのモデルバックエンドを用いた。 決定論的なフロンティア選択器の厳密なreveal正答数は、DeepSeekで96/96、GLMで82/96だった。構造化記憶を書き出す方式は、それぞれ62件成功・1件未確定と、56/96だった。設定した4結果の同時対比の有限標本識別区間は、それぞれ[0.521, 0.542]と[0.292, 0.313]であり、これは信頼区間ではない。レコード単位の監査では、元のスコアを変更せず、保持状態の十分性、応答の提供、回答スキーマへの適合を区別した。その結果、形式は正しいが意味的には誤った構造化reveal記憶が26件と25件見つかった。一方、GLMのフロンティア方式で失敗したreveal応答14件は、すべて正しい値を誤ったラッパー形式に入れていた。 削除済みであることを示すトゥームストーンを除くと、対象メカニズムで16/16件が厳密な再現に失敗した。さらに識別子の改名によって別の欠陥が明らかになり、元のフロンティア方式の後からの参照に対する十分性は、元データの8/8から変換後の94/320に低下した。ラベルの置換に対して同変な修正版を提示し検証したが、元の後からの参照に対する答えを維持できたのは2/8にとどまった。命名上の近道を除去しても、将来どの情報が関連するか不明であるという問題は解決しない。 これらの結果が裏付けるのは、適用範囲を限定した評価方法と再現可能な失敗分析であり、修正版アルゴリズムの一般的な優越性ではない。有料の予備実験からの証拠、事後診断、新たなオフライン試験は分けて報告する。独立したホールドアウト評価や、自然な実課題での検証を行ったとは主張しない。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-17(UTC)
最新改訂
2026-09-17 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

A memory can answer a current query correctly while discarding distinctions required by a later update. We investigate this failure with a paired-history audit: two histories have the same current answer, receive a shared future update, and require different subsequent answers. A pilot evaluates 24 history pairs across six synthetic mechanisms, 12 memory conditions, two repeats, and two model backends. A deterministic frontier selector obtains strict reveal accuracy of 96/96 on DeepSeek and 82/96 on GLM; a structured writer obtains 62 successes with one unresolved outcome and 56/96. The configured four-outcome joint contrast has finite-sample identification intervals of [0.521, 0.542] and [0.292, 0.313], not confidence intervals. A record-level audit distinguishes retained-state adequacy, response delivery, and answer-schema compliance without changing those original scores. It finds 26 and 25 well-formed but semantically wrong structured reveal memories, while all 14 GLM frontier reveal failures contain correct values in the wrong wrapper. Tombstone removal produces 16/16 exact replay failures in the targeted mechanism. Identifier renaming then exposes a separate flaw: original frontier late-reference adequacy falls from 8/8 to 94/320 transformed instances. We provide and test a label-equivariant repair, but it preserves only 2/8 original late-reference answers: eliminating a naming shortcut does not solve unknown future relevance. These results support a scoped evaluation methodology and reproducible failure analysis, not general superiority of the repaired algorithm. Paid pilot evidence, retrospective diagnostics, and new offline tests are reported separately; no independent held-out or natural-task validation is claimed.

著者のコメント

20 pages, 9 tables, 2 figures. Code and reproducibility materials to be released separately

arXiv ID: 2609.20045 / 要約の誤りについて