警察の事故記述を確率付きの分析変数に変換する
Calibrated Decisions at Scale: Converting Police Crash Narratives into Probabilistic Crash Variables with a System One Model (Jev)
この論文をやさしく読む
ひとことで言うと
警察の事故説明文から、定めた選択肢に該当する確率を出し、統計分析に使える項目へ変換する研究です。文章の生成は行いません。
何に役立つ?
考えられる用途は、既存の事故データの項目にない情報を抽出し、人による確認量も見積もることです。事故記述と人間の判断を使った評価を報告しています。
この研究の面白いところ
抽出精度だけでなく、確率が実際の正しさに合うかという較正と、確認すべき記録の数を扱っています。既存のコード化項目自体も、記述の正解を完全には表さないことを示します。
どこまで分かった?
再較正の改善は同じラベル上での結果です。年間10,747件は要因に帰属される事故数の増加で、実際の事故発生数の増加や、要因の因果効果の新たな実証を意味しません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
調査担当者による自由記述を含む交通事故データセットには、コード化された項目が取りこぼす情報がある。その記述の大規模なコード化には三つの障害があった。最先端の大規模言語モデルはその規模では費用が高く、生成文を検証できず、出力を人がどれだけ確認すべきかという規則もない。本論文では、記述のコード化を、条件に応じて回答する、型の定まった意思決定として定式化する。回答を担うSystem OneモデルJevは、分析者が定義した選択肢の確率を返し、文章を生成しない。テキサス州の事故記述499,500件をスクリーニングし、そのうち195,857件を27問のスキーマでコード化した。費用は記述の長さではなく、スキーマの大きさに左右される。 確率を、既存のコード化項目、および明示した標本抽出設計に基づく2,416件の盲検化された人間の判断と照合して監査した。同じ記録で最先端の大規模言語モデル二つも比較した。人間のラベルに対し、型付きモデルはF1値0.908を達成した。最先端モデルの一つはこれを0.059上回り、もう一つは型付きモデルと区別できなかった。較正の良し悪しは方式の分類よりも個々のモデルによって異なり、各モデルの監査が必要である。同じラベルを用いて再較正すると、較正誤差は3.3分の1になった。コード化項目との一致を用いると、記述への忠実度はκ係数で中央値0.26だけ過小評価された。 離散的な格子上の確率を報告する任意のモデルに適用される、分解能に起因する下限の評価を示す。また、フラグが付いた記録への確認予算により、変数ごと、年ごとに人が読む必要のある記録数を示す。較正済みの変数を既存のコード化項目に加えると、九つの要因に帰属される負傷・死亡事故数が年間10,747件増加する。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-21(UTC)
- 最新改訂
- 2026-09-21 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-21 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Crash datasets that carry an investigator narrative hold information the coded fields omit. Coding those narratives at scale has been blocked by three obstacles. Frontier large language models are costly at that scale, their generated text cannot be verified, and no rule says how much output a human must check. This paper formulates narrative coding as gated, typed decisions answered by Jev, a System One model that returns probabilities over analyst-defined options and generates no text. A screen covered 499,500 Texas narratives and 195,857 were coded with a 27-question schema. Cost is governed by schema size rather than narrative length. The probabilities are audited against coded fields and against 2,416 blinded human judgments drawn under a stated sampling design. Two frontier large language models are benchmarked on the same records. Against human labels the typed model attains an F1 of 0.908. One frontier model gains 0.059 and the other is indistinguishable from it. Calibration varies by model rather than by paradigm, so each model must be audited. Recalibration on the same labels reduces calibration error by a factor of 3.3. Agreement with coded fields understates fidelity to the narrative by a median of 0.26 in kappa. A resolution-floor bound covers any model that reports probabilities on a discrete grid. A review budget over flagged records gives the records a human must read per variable and per year. Adding the calibrated variables to the coded fields raises the injury and fatal crashes attributed to nine factors by 10,747 per year.
arXiv ID: 2609.24052 / 要約の誤りについて