文書を丸ごと保持し質問時に選ぶLLMの長期記憶
Mem++: Non-Destructive Memory for Long-Term Organizational LLM Agents
この論文をやさしく読む
ひとことで言うと
組織の文書を要約して記憶する代わりに、全文と日付・著者を保持し、質問された時点に合う文書を後から選ぶ方法です。過去の決定を上書きせずに残します。
何に役立つ?
「当時はどの決定が有効だったか」を尋ねる業務質問への回答に利用が考えられます。OrgMemBenchでは二つの回答モデルで、最良の記憶システム比較対象より8.0〜13.1ポイント高い結果を報告しています。
この研究の面白いところ
書き込み時に何を重要と判断するかを決めず、質問が来るまで情報選択を遅らせます。時間による絞り込みと字句・意味検索の融合を使い、回答モデルに版の選択を任せています。
どこまで分かった?
性能比較は挙げられたベンチマークと回答モデルでの結果です。要旨には全文保存に必要な記憶容量、検索の実行時間、アクセス権管理の評価はありません。LLM判定器による評価も含まれます。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
大規模言語モデル(LLM)のエージェントは、今や組織内の仕事に参加しており、そこでは多くの著者が数か月にわたって複数の文書へ決定事項を記録する。決定の修正は既存文書の編集ではなく新しい文書として届くため、質問に答えるには、ある時点でどの版が有効だったかを知る必要がある。しかし、ほとんどの記憶システムは書き込み時に記録を圧縮する。各文書を事実、メモ、グラフの辺へと蒸留することで、質問がなされる前に、何に答えられるかを固定してしまう。 これに対処するため、書き込み時の蒸留から読み取り時の選択へ移行する、非破壊的な記憶フレームワークMem++を提案する。Mem++は全ての文書を日付と著者とともに丸ごと保存し、書き込み時には生成モデルを呼び出さない。読み取り時には、質問が対象とする時点以前の日付の文書だけを取得し、字句的な順位付けと意味的な順位付けを融合する。古い版を上書きするシステムとは異なり、Mem++は古い版を残し、選択を回答モデルに委ねる。 組織向けベンチマークOrgMemBenchでの評価では、二つの回答モデルを通じて、Mem++は最も強い記憶システムのベースラインを8.0〜13.1ポイント上回る。gpt-4.1-miniでは、RAGを2.6ポイント上回る最高の総合スコアも達成する。さらに、LoCoMoではLLM判定器の平均スコアが最高となり、LongMemEval-Sでは自身のエンティティグラフ版に次ぐ2位となる。ベンチマーク評価用コードは https://github.com/AIDAChip-Inc/mem-plus-plus で公開している。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-10-01(UTC)
- 最新改訂
- 2026-10-01 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-10-01 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Large Language Model (LLM) agents now take part in organizational work, where many authors record decisions across documents over months. Because a revised decision arrives as a new document rather than an edit, answering a question requires knowing which version held at a given time. However, most memory systems compress the record at write time. By distilling each document into facts, notes or graph edges, these methods fix what can be answered before any question is asked. To address this, we propose Mem++, a non-destructive memory framework shifting from write-time distillation to read-time selection. Mem++ stores every document whole with its date and author, and it calls no generative model at write time. At read time, it retrieves only documents dated up to the time a question asks about and fuses lexical and semantic rankings. Unlike systems that overwrite older versions, Mem++ keeps them and leaves the choice to the answering model. Evaluations on the organizational benchmark OrgMemBench demonstrate that Mem++ surpasses the strongest memory system baseline by 8.0 to 13.1 points across two answering models. With gpt-4.1-mini, it also achieves the best overall score, 2.6 points above RAG. In addition, Mem++ achieves the best average LLM-judge score on LoCoMo and ranks second on LongMemEval-S, behind only its entity-graph variant. Code for benchmark evaluation is available at https://github.com/AIDAChip-Inc/mem-plus-plus.
著者のコメント
15 pages, 4 figures
arXiv ID: 2610.02002 / 要約の誤りについて