arXiv論文メモ
新着一覧
stat.ML / cs.LG · 査読状況未確認

編集や混入に強いLLM生成文の統計的検出

Robust Detection of LLM-Generated Text under Contamination

Jiaxun Li, Saptarshi Chakraborty, Ambuj Tewari

この論文をやさしく読む

ひとことで言うと

文章が人間とLLMのどちらから来たかを、編集や別の文章の混入があっても判定しやすくする統計手法です。

何に役立つ?

生成文検出器の頑健性を評価し、尤度比などの極端な値を切り詰める設計に役立ちます。要旨では七つの検出器と複数のデータセットで比較しています。

この研究の面白いところ

ある量以上の混入では検出そのものが不可能になる境界を示し、その下での切り詰め検定の性質を証明しています。

どこまで分かった?

理論は有限次数Markov過程とHuber型混入の仮定の下で成り立ちます。実験の改善幅は検出器と混入条件で変動します。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

編集や混入を受けた文章から、LLM生成文を検出する問題を研究する。人間と機械の文章をHuber型の混入を伴う有限次数Markov過程でモデル化し、この仮定の下で信頼できる検出が可能になる厳密な境界を特徴付ける。混入の大きさが、混入のない二つの生成源の隔たりに比べて十分大きい場合、検出は不可能である。この境界より下では、尤度比を切り詰めた複数の検定により、最悪の場合の誤りをゼロへ近づけられる。この構成は、既存の統計的検出器に切り詰めを加えるという単純な変更を動機付ける。広い範囲の加法的スコアについて、切り詰めた検定は一致性を持つ一方、切り詰めない検定の最悪条件での検出力がゼロに近づく条件を特定する。 三つのデータセットと三つの生成モデル、およびRAIDベンチマークで七つの検出器を評価したところ、切り詰めにより両方の研究で頑健性が改善したが、その利得は検出器や混入条件によって異なった。例えば目標の偽陽性率を5%とすると、対数尤度と対数ランクの比に基づくLRR検出器の真陽性率は、制御された研究で中央値8.3パーセントポイント、RAIDの割合別評価と攻撃別評価でそれぞれ2.1ポイントと4.3ポイント改善した。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-24(UTC)
最新改訂
2026-09-24 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

We study the detection of LLM-generated text under editing and contamination. Modeling human and machine text as finite-order Markov processes with Huber contamination, we characterize an exact boundary for reliable detection under our assumptions. Detection is impossible when contamination is sufficiently large relative to clean-source separation. Below this boundary, a collection of clipped likelihood-ratio tests achieves vanishing worst-case errors. This construction motivates clipping as a simple modification of existing statistical detectors. For a broad class of additive scores, we identify conditions under which the clipped test is consistent while the raw test's worst-case power tends to zero. We evaluate seven detectors across three datasets and three generation models, and on the RAID benchmark. Clipping improves robustness in both studies, with gains varying across detectors and contamination settings. For example, at a target false-positive rate of 5\%, clipping improves the log-likelihood--log-rank ratio (LRR) detector's true-positive rate by a median of 8.3 percentage points in the controlled study and 2.1 and 4.3 points in rate- and attack-specific RAID evaluations, respectively.

arXiv ID: 2609.29935 / 要約の誤りについて