企業や世代をまたいで言語モデルの振る舞いを低コストで測る方法
Low-Cost Assays for Measuring Model Behavior Across Vendors and Releases
この論文をやさしく読む
ひとことで言うと
同じ公開テストをさまざまな言語モデルに繰り返し実施し、応答と実際の行動を比較する低コストの測定方法を示した研究。
何に役立つ?
モデルの更新や開発元による振る舞いの違いを継続的に調べる際の測定設計に役立つ。論文が示したのは特定のテストでの比較結果であり、あらゆる場面での安全性評価ではない。
この研究の面白いところ
答えの語句、誘導的な文末表現への反応、圧力下での立場、指示に反する文書があるときの実際の行動を、それぞれ異なる測定法で扱っている。
どこまで分かった?
結果は提示した固定の刺激と実行環境に基づく。同じモデルでも実行環境で行動が変わり得ることを要旨が示しており、単一の測定値だけでモデル全体の性質を決められるとは述べていない。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
言語モデルは助言や対話を行い、人が眠っている間にもソフトウェアを書く。その振る舞いを測るのは難しい。モデル、プロンプト、リリースをまたいで繰り返しサンプルを取り、主に自由形式の文章で得られる応答を集計前に分類し、企業間やモデル間で意味のある比較ができるよう、結果を分かりやすく厳密に示す必要があるからだ。これらの制約に対応するため、本研究では、モデルの振る舞いを調べる簡単で安価、拡張可能かつ再現可能な方法を提示する。各調査では、固定して公開した刺激を複数企業のモデル群に同一条件で与え、費用はモデル当たり数ドル以下に抑える。得られた対話記録は、必要な解釈の程度に応じて三つの方法のいずれかで読む。応答形式を制約した場合の完全一致判定、コードごとに人間の分類者との一致度を報告するLLM判定者による分類基準の適用、またはエージェントが発言とは別に実際に何をしたかを記録する計測環境である。 先端モデルとオープンソースモデルの研究所による4年間のリリースに適用すると、四つの結果が得られた。収束:単語を一つ選ばせると、44モデル中27モデルが4回の試行のうち少なくとも1回は「serendipity」と答えた。抵抗:「right?」を文末に付けると賛同率が最大32ポイント変化し、その表面的な言い回しに応じて、世代が進むにつれ迎合的な反応から抵抗的な反応へ方向が反転した。立場:圧力を受けても立場を維持するかどうかはモデルの世代と関連し、どのように維持するかは開発元の研究所と関連した。説明責任:リポジトリの文書に反することを指示されたとき、黙って従うことが一度もないコーディングエージェントもあれば、常に従うものもあり、同じモデルでも実行環境によって変わり得た。こうした一連のテストをリリースごとに繰り返せば、企業や時間をまたぐ振る舞いの変化を追跡できる。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-24(UTC)
- 最新改訂
- 2026-09-24 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-24 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Language models advise people, keep them company, and write software while they sleep. Measuring what they do is hard: behavior has to be sampled repeatedly across models, prompts and releases, most of it lives in unstructured text that has to be coded before it can be counted, and the result has to be legible and rigorous enough to meaningfully compare models and vendors. To address these constraints, we present a simple, cheap, scalable, and replicable model for studying model behavior. Each study is a frozen, public stimulus run identically on a cross-vendor panel, at a few dollars per model or less. Each reads its transcripts one of three ways, chosen by how much interpretation the behavior needs: exact match on a clamped reply, a codebook applied by LLM judges whose agreement with a human coder is reported per code, and an instrumented environment that records what an agent did independently of what it said. Run across four years of model releases from both frontier and open-source labs, these instruments find four things. Convergence: asked to pick a word, 27 of 44 models answer serendipity at least once in four tries. Resistance: a trailing "right?" moves endorsement by up to 32 points, and the sign flips from sycophantic to resistant as generations advance, keyed to the tag's surface form. House: whether a model holds a position under pressure tracks its generation, and how it holds tracks the lab that built it. Account: told to do something the documentation in their repository contradicts, some coding agents never went along silently and others always did, and the same model can change with the harness it runs in. Re-run on every release, batteries like these track how behavior is changing across vendors and over time.
著者のコメント
6 pages. Code and data: https://github.com/tap2k/modelun
arXiv ID: 2609.30012 / 要約の誤りについて