医療の系統的レビューで根拠を確認できるAI抽出支援
EviStreams: Human-in-the-Loop AI Data Extraction for Systematic Reviews in Medicine
この論文をやさしく読む
ひとことで言うと
医療の系統的レビューで、二重確認と根拠の記録を保ちながらAI抽出を使うウェブツールである。
何に役立つ?
考えられる用途は、レビュー班のデータ抽出作業の支援である。要旨は人による確認と判定を含む作業手順を示している。
この研究の面白いところ
プロンプトではなく型付きの項目を専門家が定義し、抽出値を根拠の文章と並べて確認できる。評価では項目定義の影響がモデル選択より大きかった。
どこまで分かった?
評価は4つの臨床コーパスと3種類のモデル群で行われた。どのレビューにも同じ精度で適用できるとは要旨に示されていない。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
系統的レビューは診療ガイドラインの基礎になるが、データ抽出には多くの専門家の労力がかかる。規定された作業手順では、二人の査読者が各研究から独立にデータを抽出し、判定担当者が不一致を解決し、各値がどう得られたか監査可能な記録を残す必要がある。大規模言語モデルは抽出を支援できるが、既存のレビュー手順に適合し、再現性を保たなければならない。 本論文は、レビュー班がAI支援の抽出を三つの重要な段階で管理できる、公開中のオープンソースでコード不要のウェブプラットフォームEviStreamsを提示する。三段階は、コードを動かす前に承認する構造化されたプログラム設計、試行で調整する型付きの項目定義、査読者間で互いの結果を見せない二重レビューと判定を伴う抽出予測である。専門家はフォーム作成機能を使い、プロンプトではなく型のある項目を定義し、アップロードしたPDFからデータを抽出する。各値を根拠となる文章と並べて確認し、互いの結果を伏せた二重レビューの不一致を解消して、監査可能な合意結果を出力する。 システムとともに公開された、臨床分野の4つのコーパスと3種類の先端モデル群を使う評価では、抽出の品質はモデルの選択よりも項目の定義に強く左右された。EviStreamsは https://evistreams.com/demo で公開され、Apache-2.0ライセンスで提供されている。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-23(UTC)
- 最新改訂
- 2026-09-23 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-23 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Systematic reviews underpin clinical guidelines, yet their data-extraction step is a major expert-labor bottleneck bound by a protocolized workflow: two reviewers extract each study independently, an adjudicator resolves disagreements, and the team keeps an auditable record of how every value was produced. Large language models can assist with extraction, but that assistance must fit established review protocols and preserve reproducibility. We present EviStreams, a live, open-source, no-code web platform that puts review teams in control of AI-assisted extraction at three key stages: program design (a structured decomposition approved before any code runs), field specification (typed field definitions calibrated from a pilot), and extracted predictions (reviewer-blinded dual review with adjudication). Working through a form builder, a domain expert defines typed fields rather than prompts, runs extraction over uploaded PDFs, inspects every value alongside the supporting passage it came from, and resolves a reviewer-blinded dual review into an auditable consensus export. An evaluation across four clinical corpora and three frontier model families, released with the system, shows that extraction quality is shaped far more by the field specification than by the choice of model. EviStreams is live at https://evistreams.com/demo and released under Apache-2.0.
著者のコメント
12 pages, 5 figures. Accepted to EMNLP 2026 System Demonstrations
arXiv ID: 2609.27418 / 要約の誤りについて