製品・政策変更への反応を模擬するAugurの評価
Augur: A Synthetic Decision Lab for Rehearsing Reactions to Product and Policy Changes
この論文をやさしく読む
ひとことで言うと
製品や政策の変更に対する人々の反応を模擬する仕組みと、その評価方法を調べた研究です。
何に役立つ?
変更前の懸念を洗い出す用途が考えられます。実証結果としては、実際に挙がった懸念の67~90%を合成反応が再現したという評価が示されています。公開判断の正確さとは別の指標です。
この研究の面白いところ
プロンプトで判断の分類を明示するだけで得点が大きく変わり、モデル間の見かけ上の能力差の多くが評価条件に左右されました。
どこまで分かった?
Gold-50は実際の結果が分かる50事例です。全体の処理系には過度に悲観的な判断の偏りがあり、反応の再現率が高くても公開判断の正確さを保証しません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
製品や政策の変更を実施する前には、人々がどう反応するかが重要となる。Augurは、その反応をオフラインで予行演習する。変更文書から型付きの知識グラフを構築し、根拠付けられた人物像の集団を配置して相互作用をシミュレーションし、五つの行動のいずれかを勧める監査可能な意思決定メモを返す。著者らは、実際の結果が分かり、公的記録と照合された製品・政策の50事例からGold-50を作り、五択の公開判断を評価した。 中心的な発見は方法論上の否定的なものである。最先端のクラウドモデルと、追加学習してオフラインで動かす公開重みモデルとの測定上の差の大半は、能力差ではなく評価条件の指定不足に起因した。これを三つの形で示す。第一に、重み、事例、採点方法を固定しても、入力文を包むプロンプトだけで得点が大きく変わり、Qwen3-32BのLoRA-SFTアダプターは0%から73%まで変動した。第二に、条件をそろえた2×2の比較では、判断の分類をプロンプトに定義するだけで、モデルを変えずに各最先端モデルの得点が24~34パーセントポイント上がった。指定不足のプロンプトでは、オフラインで動くQwen3-32B LoRA-SFTが三つの最先端モデルすべてを上回ったが、プロンプトを公平にすると、いずれとも有意差は検出されなかった。比較には対応のあるMcNemar検定とHolm補正を用いた。第三に、蒸留元モデルとの一致率が上がっても正解率は追随せず、全体の処理系は判断を過度に悲観する系統的な偏りを改善せず増幅した。 反応生成部分は別に検証した。四つのモデル群からなる盲検の判定者は、合成反応が実際に公衆から挙がった懸念の67~90%を再現すると評価した。事前登録した要素除去実験では、この部分の価値は判断が最も難しい場合に大きく、成績が上限に近い場合には重複すると分かった。著者らは、ここに示した数値と図をすべて再生成する処理系を公開している。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-24(UTC)
- 最新改訂
- 2026-09-24 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-24 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Before a product or policy change ships, the question that matters is how people will react to it. Augur rehearses that reaction offline: it builds a typed knowledge graph from the change documents, populates a grounded persona market, simulates the interaction, and returns an auditable decision memo recommending one of five actions. We assemble Gold-50, fifty real product and policy episodes whose real-world outcome is known, adjudicated against the public record, and score the five-way release verdict against it. Our central finding is methodological and negative: most of the measured gap between frontier cloud models and open-weight models we fine-tune and serve offline is attributable to an under-specified evaluation, not a difference in capability. We show this three ways. First, the prompt envelope alone can dominate the score: holding weights, cases and scorer fixed, one system -- a LoRA-SFT adapter on Qwen3-32B -- swings from 0% to 73%. Second, in a matched 2x2 ablation, defining the decision taxonomy in the prompt -- with no model change -- lifts every frontier model by +24 to +34pp; under the under-specified prompt, Qwen3-32B LoRA-SFT served offline beats all three frontier models (paired McNemar, Holm-corrected), and once the prompt is fair no significant difference from any of them is detected. Third, agreement with the distillation teacher rises without accuracy following, and the full pipeline amplifies a systematic "over-doom" bias rather than improving the verdict. Separately, we validate the reaction layer on its own terms: blind judges across four model families find the synthetic reaction recovers 67-90% of the concerns the public actually raised, and a pre-registered ablation locates its value -- largest where the decision is hardest, redundant near ceiling. The pipeline that regenerates every number and figure here is available from the authors.
著者のコメント
19 pages, 15 figures, 11 tables
arXiv ID: 2609.29952 / 要約の誤りについて