言語モデルが19~20世紀初頭の英語を再現できるか測る
Chronologic: Measuring Language Models' Ability to Represent the Past
この論文をやさしく読む
ひとことで言うと
1831~1930年の英語の文脈に、言語モデルの回答がどれほど合うかを歴史資料で測る研究。
何に役立つ?
歴史研究でモデルの回答を参考にする際、時代に合う表現かを点検する方法の検討に役立つ。要旨では、現行モデルの自由生成には課題が残ると報告される。
この研究の面白いところ
複数の正答と強い紛らわしさを持つ不正解を用い、単一正答を前提としない評価を設計した。モデルが自分の生成文の弱さを見分けられる点も興味深い。
どこまで分かった?
対象は1831~1930年の英語圏の文脈であり、他の時代や言語での性能は要旨からは分からない。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
言語モデルは過去を研究する道具として魅力的である。しかし、モデルが示す証拠を信頼するには、その応答が対象時代に合っているかを研究者が知る必要がある。検証は難しい。現代に生きる人が通常行わない課題であり、多くの問いには正しい答えが複数あるためである。 著者らは歴史的なテキストを使い、1831~1930年の英語圏の文脈をモデルがどのように表現できるかを測るベンチマークを作成した。最も難しい問いを適切に段階的に採点するため、複数の正解と紛らわしい不正解とのペア比較に依拠する。結果として、文章を生成する課題は、選択肢を見分ける課題より難しかった。実際、推論モデルは、自身が生成した答えの弱点を通常は見抜けた。 歴史的な文章だけで事前学習したモデルは、正解の尤度で評価すると先行する一方、自由生成では商用モデルに及ばなかった。調べたモデルの中に、歴史的な文脈を十分に説得力を持って表現できるものはまだなかったが、その目標に向けた進展は見られた。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-19(UTC)
- 最新改訂
- 2026-09-19 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-19 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Language models are appealing tools for research on the past. But to trust the evidence a model provides, researchers need to know whether its responses fit the period represented. Validation is challenging, because this is not a task living people ordinarily perform, and because many questions have multiple correct answers. We use historical texts to develop a benchmark for a model's representation of English-language contexts 1831-1930, relying on pairwise comparisons to multiple ground truths and strong distractors to score the hardest questions in an appropriately graduated way. We find that generative tasks are harder than discriminative ones; in fact, reasoning models can typically discern the weakness of their own generated answers. While models pretrained exclusively on historical text lead the pack when evaluated by answer likelihood, they cannot compete with commercial models in free generation. None of the models we tested represent historical contexts in a fully persuasive way yet, but progress toward that goal is evident.
著者のコメント
22 pages, 3 figures, 8 tables
arXiv ID: 2609.23178 / 要約の誤りについて