レイアウト領域のマスクでGROBIDの学術PDF解析を改善
Layout-Guided Masking for GROBID: Lightweight Structural Gains in Large-Scale Scientific PDF Ingestion
この論文をやさしく読む
ひとことで言うと
図や表などの領域をCPUで検出してGROBIDに渡し、学術PDFの構造抽出を改善した。
何に役立つ?
大量の学術PDFを費用を抑えて機械可読化する際、CPUで動く構造解析の選択肢になる。
この研究の面白いところ
段落や節だけでなく表検出も大きく改善し、費用をGPU方式と比較している。
どこまで分かった?
表構造では最も強いGPU方式に届かず、結果は示されたPMC二分野と外部表ベンチマークでの評価である。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
学術PDFを機械可読な全文に変換する作業は、大規模情報システムの障害となっている。近年の画像認識型パーサーは精度を改善する一方、GPUを要し、抽出した文章にノイズを加えることがある。CPUで動くフォントストリーム型のモジュール式パーサーGROBIDは、学術論文の構造化に広く使われ、大規模な公開学術コーパスを支えている。本研究では、図、表、欄外要素(ヘッダー、フッター、ページ番号)の領域を特定する軽量なCPU検出器を組み合わせる。検出領域を種類付きの領域マスクとして表し、その中のトークンをGROBIDの専用モデルへ振り分けるか破棄する。PMCの生命情報学1,926件と材料科学2,595件の二つのコーパスで、JATSを基準に節を考慮した構造評価を行った。拡張版は素のGROBIDよりほとんどの指標で改善し、NSはそれぞれ+0.025、+0.013、材料科学の段落再現率は+0.086(d_z=1.08)となり、キャプションと対応づいた図の回収率も両コーパスで向上した。外部のTable-BRGMベンチマークでは、表検出のF1は0.16から0.94に、表構造のGriTS-Topは0.27から0.78に改善したが、最も強いGPUシステムには及ばなかった。本文については、画像認識型の四システム(Docling、MinerU、olmOCR、dots.ocr)との比較で、両コーパスで最高の段落適合率、材料科学で最高の節検出を示し、文字誤り率は最良のGPUパーサーとの差が0.004以内だった。CPUでの一連の処理の費用は、最も安価なGPUシステムDoclingの2.7~3.2分の1、生成型パーサーの10~14分の1だった。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-22(UTC)
- 最新改訂
- 2026-09-22 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-22 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Transforming scholarly PDFs into machine-readable fulltext remains a bottleneck for large-scale information systems. Recent vision-based parsers improve accuracy, but need GPUs and may introduce noise into the extracted text. GROBID, a modular font-stream parser running on CPU, is the de-facto standard for structuring scientific articles and underpins several of the largest open scholarly corpora. We pair it with a lightweight CPU detector localising figure, table, and paratext (header, footer, page number) regions, encoded as typed-area masks whose tokens are routed to GROBID's specialised models or discarded. On two PMC corpora, Bioinformatics (1,926 articles) and Materials Science (2,595), scored against JATS with a section-aware structural protocol, our extension improves over plain GROBID on most metrics (NS $+0.025$/$+0.013$; $+0.086$ paragraph recall on Materials Science, $d_z{=}1.08$), and caption-linked figure recovery improves on both corpora. On the external Table-BRGM benchmark, table detection recovers F1 $0.16 \to 0.94$ and table structure follows (GriTS-Top $0.27 \to 0.78$, below the strongest GPU system). On body text, against four vision-based systems (Docling, MinerU, olmOCR, dots.ocr), it has the best paragraph precision on both corpora, the best section detection on Materials Science, and a character error rate within 0.004 of the best GPU parser. End-to-end on CPU, it costs $2.7$--$3.2\times$ less than the cheapest GPU system (Docling) and $10$--$14\times$ less than generative parsers.
arXiv ID: 2609.26381 / 要約の誤りについて