LLMの汎化を正答率だけでなく応答の安定性から評価
Generalization Is Stability, Not Accuracy: Multi-Axis Evaluation of LLMs
この論文をやさしく読む
ひとことで言うと
同じ内容を言い換えたときの答えや内部状態の変わり方を、複数の軸から測るLLM評価です。
何に役立つ?
平均正答率だけでは見えにくい、表現の違いに対する不安定さを調べるのに役立ちます。
この研究の面白いところ
出力の一致だけでなく、内部活性や確信度なども測り、一つの得点では異なる失敗を隠してしまう可能性を検討しています。
どこまで分かった?
要旨にはモデル数や具体的な統計量が示されていません。また、応答のミラーリングの操作的な定義は要旨だけでは分かりません。結果は評価対象モデルとデータセットについてのものです。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
大規模言語モデル(LLM)の汎化とは、同じ入力を異なる形で表現したときにも、一貫し、意味的に安定した出力を生成する能力である。既存研究は通常、一つのプロンプト形式、課題、または変形の集合に対する集約正答率で汎化を評価するため、頑健性とベンチマーク全体の性能を混同してしまう。 本研究では、個々の事例のレベルで、複数の入力変形とモデル挙動の異なる側面をまたいで汎化を評価する。限定的な訓練や、汎化の評価を不透明にするほかの手段によって改善できる一つの得点に性能を還元するのではなく、変動に注目する。この考え方に従い、安定性を考慮した汎化目的SAGOを導入する。これは、同じ入力に対して、変形やベンチマークが異なるとモデルの挙動がどれだけ変わるかを測る枠組みである。生成の一貫性、内部活性、確信度、応答のミラーリングなど、複数の次元にわたる変動を捉える。 広く使われる多くのモデルが、統計的に有意で一貫した汎化の不安定性を示すことを明らかにする。どのモデルも一様には汎化せず、挙動の各軸は独立した失敗様式を捉え、データセットを変えることでモデル順位が逆転することもある。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-10-01(UTC)
- 最新改訂
- 2026-10-01 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-10-01 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Generalization in large language models (LLMs) is the ability to produce consistent and semantically stable outputs when the same input is expressed in different ways. Existing work typically evaluates generalization through aggregate accuracy on a single prompt format, task, or set of variations, which conflates robustness with overall benchmark performance. In this work, we show generalization evaluation at the level of individual examples, across multiple input variants, and across different aspects of model behavior, focusing on variability rather than reducing performance to a score that can be improved through narrow training or other ways that obfuscate generalization evaluation. Following this view, we introduce the Stability-Aware Generalization Objective (SAGO), a framework that measures how much model behavior changes for the same input under different variations and benchmarks, capturing variability across several dimensions including generation consistency, internal activations, confidence, and response mirroring. We show that many commonly used models exhibit statistically significant and consistent generalization instability: no model generalizes uniformly, behavioral axes capture independent failure modes, and cross-dataset variation can reverse model rankings.
著者のコメント
Accepted at the TAE (Trust-AI-Eval) Workshop: Can We Trust AI Evaluation?, NeurIPS 2026
arXiv ID: 2610.01428 / 要約の誤りについて