偽の算数事実で学習した言語モデルの自信を測る手順書
Technical Manual for Toolkit for Confidence-Corpus Consistency via Fine-Tuning on a Fabricated Corpus
この論文をやさしく読む
ひとことで言うと
言語モデルが自信を持って答えることと、正しい事実を知っていることの関係を調べる実験手順書。
何に役立つ?
偽の情報で追加学習したモデルの回答確信度を、測定のずれを抑えて比較する実験を組む際の参考になる。
この研究の面白いところ
81通りの足し算に一貫した偽の答えを与え、学習前後で同じ測定法を使う。答えの桁数によるトークン化の違いも考慮する。
どこまで分かった?
この手順書は器具の説明であり、特定の実験の結果や解釈を報告していない。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
言語モデルが回答に示す自信は、その事実をどれだけ知っているかの代わりに読まれがちである。本手順書は、この解釈を直接調べるために作られた公開ツールキットを説明する。小型の因果言語モデルを、1桁同士の足し算81通りのそれぞれについて一つの偽の答えを一貫して主張する文章群で追加学習する。その後、各偽の答えに対するモデルの自信を、追加学習前に対応する正しい答えに示した自信と比較する。測定手順は前後で変えない。 事実の組合せの生成、トークン長を考慮した自信の測定、基準値の検証、文章群の構築、追加学習、学習前後の対応を取った比較という、作業の各段階を説明して根拠を示す。また各段階が排除しようとする交絡要因も説明する。たとえば、1桁と2桁の答えでトークン分割が非対称になることや、ある答えが単に優位性を失った場合と積極的に抑制された場合の違いである。この原稿は手法と実装の参照資料であり、測定器具を記録するもので、特定の実行結果を報告・解釈しない。ツールキットと依存関係を固定した環境は、永続識別子の下で別途保存されている(第9節)。この器具で実験結果を得て解釈する研究では、その保存資料を引用できる。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-23(UTC)
- 最新改訂
- 2026-09-23 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-23 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
A language model's confidence in an answer is often read as a proxy for how well it knows the corresponding fact. This manual documents an open toolkit built to test that reading directly: a small causal language model is fine-tuned on a corpus that consistently asserts one fabricated arithmetic answer for each of the 81 single-digit addition pairs, and its post-fine-tuning confidence in each fabricated answer is compared against its own pre-fine-tuning confidence in the corresponding true answer, using an unchanged measurement procedure throughout. We describe and justify every pipeline stage, fact-space generation, token-length-aware confidence measurement, baseline validation, corpus construction, fine-tuning, and paired before/after comparison, together with the confound each is meant to rule out, among them tokenization asymmetry between single- and double-digit answers and the difference between an answer merely losing its edge and one being actively suppressed. This manuscript is a methodological and implementation reference: it documents the instrument and does not report or interpret the outcome of any specific run. The toolkit and its pinned dependency environment are archived separately (Section 9) under a persistent identifier, to be cited as an instrument by work that produces and interprets empirical results with it.
著者のコメント
30 pages, 2 figures, 1 table, 12 code listings. Methodological and implementation reference manual; does not report or interpret empirical results from any specific run. Toolkit and pinned dependency environment archived at https://doi.org/10.5281/zenodo.22903853 (CC BY 4.0)
arXiv ID: 2609.28747 / 要約の誤りについて