arXiv論文メモ
新着一覧
cs.AI · 査読状況未確認

別のLLMの内部表現から誤情報の始まりを検出する

External Observers May See More Clearly: Cross-Model Span-Level Hallucination Detection in Large Language Models via Hidden State Probing

Kingshuk Gupta, Davide Buscaldi

この論文をやさしく読む

ひとことで言うと

LLMが作った文章のどこから誤情報が始まるかを、層ごとの内部表現を使って検出する方法です。

何に役立つ?

考えられる用途は、長い生成文のうち確認が必要な範囲を絞り込むことです。文章全体を正誤判定するだけでなく、開始点と継続部分に注目します。

この研究の面白いところ

生成したモデル自身より、小さい別の観察モデルの方が開始点をよく検出できる場合があります。

どこまで分かった?

要旨には具体的なAUC、モデル名、データセット名は示されていません。比較されるのは誤情報の開始位置の検出であり、文章全体の事実性を保証したり、外部検索の必要性をなくしたりする結果ではありません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

大規模言語モデル(LLM)が推論の基盤としてますます使われる一方、誤情報を生成するハルシネーションの傾向は重大な弱点であり続ける。近年の内部状態プローブは、時間のかかる外部検索システムに代わる有望な方法だが、多くはハルシネーション検出をトークンごとの二値分類に縮約しており、意味の逸脱が持つ構造化された逐次的な境界を捉えられない。 本研究では、細かなスパン単位でハルシネーションを検出する、内部隠れ状態に基づく枠組みを導入する。層ごとの活性化パターンを調べ、LLMの生成中にハルシネーションが正確に始まるトークンと、その継続トークンを検出しようとする。実験は、この方法がハルシネーションの開始をうまく切り分け、極端なクラス不均衡にもかかわらず、ランダムな比較基準より適合率・再現率曲線下面積を大きく改善することを示す。 最終的に、一つのモデルが別のモデルの生成によって誘発される内部表現を観察する、新しいモデル間検出の枠組みを提案する。外部の観察モデルは、生成モデルによる自身のハルシネーション開始の検出に匹敵し、または上回れることが分かる。観察モデルの方が小さい場合も含まれ、開始位置の特定において、自己検出が性能の上限ではないことを示唆する。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-10-01(UTC)
最新改訂
2026-10-01 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

As Large Language Models (LLMs) increasingly serve as foundational reasoning engines, their tendency to hallucinate remains a critical vulnerability. While recent internal state probes offer a promising alternative to slow external retrieval systems, they largely reduce hallucination detection to a token-wise binary classification task, failing to capture the structured, sequential boundaries of semantic drift. Here, we introduce an internal hidden state framework for fine-grained, span-level hallucination detection. By inspecting layer-wise activation patterns, we attempt to detect the exact hallucination onset and continuation tokens in an LLM generation. Our experiments show that this approach successfully isolates hallucination onsets, achieving substantial improvements in Precision-Recall AUC over random baselines despite extreme class imbalance. Ultimately, we propose a novel cross-model detection framework in which one model observes the internal representations elicited by another model's generation. We find that an external observer can match or exceed a generator's self-detection of its own hallucination onsets, including when the observer is the smaller model, suggesting that self-detection is not the ceiling for onset localisation.

著者のコメント

12 pages, 2 figures, 9 tables

arXiv ID: 2610.02066 / 要約の誤りについて