医師が確認した架空の電子カルテ評価データ
Synthetic Hospital: An Open, Verifiable, Physician-Validated Longitudinal EHR Benchmark
この論文をやさしく読む
ひとことで言うと
公開可能で正解を検証できる架空の電子カルテを作り、AIの長期的な患者記録の理解を評価した。
何に役立つ?
医療情報を公開せずに、カルテの経時的な理解や重要所見の要約を評価するデータとして使える。
この研究の面白いところ
医師が実記録と区別する成績はほぼ偶然水準だった一方、最良のAIでも重要所見を約半分見落とした。
どこまで分かった?
完全に合成した患者記録であり、実際の臨床現場での診断や安全性を直接検証したものではない。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
先端的な言語モデルが臨床業務であまり使われない一因は、開発に必要な現実的で長期にわたる評価用データが少ないことである。実際の電子カルテはプライバシー、倫理、利用条件のため公開しにくく、記録は医療者が書いた内容に限られるため検証可能な正解も備えていない。本研究は、公開できることと検証可能な正解を持つことの両方を目指し、完全に合成した事実に基づく長期的な電子カルテのベンチマークSynthetic Hospitalを示す。保護対象の医療情報を使わず、公開された医学教育資料だけから作り、長期的に追跡する1,268人の患者と5,602回の診療を含む。すべての診断、所見、時間的関係には、ICD-10-CM、SNOMED CT、LOINCなどの標準用語体系による裏付けと、元の教材までたどれる完全な出典の連鎖がある。 Synthetic Hospitalは、標準的な相互運用API、役割に基づくアクセス、関数呼び出しのインターフェースを備え、実際の電子カルテ基盤を模した病院記録システムから提供される。医師による盲検評価では、合成記録と実際の患者記録を区別する正解率は偶然に近い53%だった。先端モデルと公開モデル計10種類の評価で、上限に近づくモデルはなかった。最良モデルが患者の長期的な問題一覧を再構成した際の重症度で重み付けしたF1は0.73で、条件をそろえた一部データに対する医師7人の平均と同じだが、最良の医師の0.89を大きく下回った。またカルテの要約では臨床上重要な所見のおよそ半分を見落とした。結果は、Synthetic Hospitalが臨床AIの性能を試す難しく現実的な評価課題であることを示す。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-24(UTC)
- 最新改訂
- 2026-09-24 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-24 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Frontier language models are rarely used in clinical workflows because the realistic, longitudinal benchmarks needed to develop them are scarce. Real electronic health record (EHR) data cannot be openly shared due to privacy, ethics or data use issues and it does not contain verifiable ground truth since the chart records only reflect what clinicians documented. We introduce Synthetic Hospital, an open, fully synthetic, fact-grounded longitudinal EHR benchmark that resolves the open sharing and verifiable ground truth barriers. Built entirely from public medical-education material with no protected health information, it comprises 1,268 longitudinal patients and 5,602 encounters, where every diagnosis, finding, and temporal relation is grounded in standard ontologies (ICD-10-CM, SNOMED CT, LOINC) and with a complete provenance chain back to its source medical education material. Synthetic Hospital is served through a simulated hospital record system that mirrors real EHR infrastructure (standard interoperability APIs, role-based access and function-calling interface). In a blinded review, physicians distinguished its records from real patient charts at near-chance rates (53\%). Across 10 frontier and open models, none approaches ceiling: the best model reconstructs a patient's longitudinal problem list with a severity-weighted F1 of 0.73, level with the mean of seven physicians on a matched subset but well below the best of them (0.89), and misses roughly half of clinically relevant findings when summarizing a chart. Overall, these results highlight that Synthetic Hospital is a difficult and realistic test of clinical AI performance.
著者のコメント
29 pages, 2 figures, 12 tables
arXiv ID: 2609.30027 / 要約の誤りについて