arXiv論文メモ
新着一覧
cs.CV · 査読状況未確認

文書解析の弱点に合わせて学習データを改善するWeVisDoc

WeVisDoc: From Coverage to Capability for Robust End-to-End Document Parsing

Hao Yu, Kang Liu, Linnan Zhao, Jiabo Zhan, Chong Sun, Chen Li, Jing Lyu

この論文をやさしく読む

ひとことで言うと

文書画像を読み取るAIの弱点を測り、その弱点に合わせて学習データを補う方法です。

何に役立つ?

レイアウトや撮影品質がさまざまな文書を構造化する作業に役立ちます。単にデータの種類を増やすだけでなく、残る誤りに学習資源を配分します。

この研究の面白いところ

第1段階で対象の幅を広げ、第2段階で視覚・構造のグループ別に誤りを診断します。4Bモデルでは実際に劣化した文書の評価が第1段階から4.03点改善しました。

どこまで分かった?

OmniDocBench v1.6で95.38、PureDocBenchの3部門平均で75.54を報告しています。首位は比較対象に含まれる端から端まで処理するパーサーの範囲で、すべての文書形式への保証ではありません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

文書解析は文書画像を構造化された内容へ変換するものであり、多様なレイアウトや画像取得条件のもとで安定した性能が求められる。しかし、学習コーパスは一般的な文書の種類やきれいなデジタルページに偏っている。また、対象範囲を広げるだけでは、解析器に残る弱点への対処方法は定まらない。 頑健なエンドツーエンド文書解析のために、データを中心とする2段階の枠組みWeVisDocを提案する。第I段階では、異種データと、構造を保った劣化画像の合成によって、意味・構造・外観の対象範囲を広げる。第II段階では、学習から分離した評価用データを使い、固定した視覚・構造クラスタ内で、第I段階の解析器に残る誤りを測定する。この診断結果に基づいて、弱点に狙いを定めたデータ構築と、出力対象トークンの予算の再配分を行う。 WeVisDoc-4Bは、OmniDocBench v1.6でOverallスコア95.38、PureDocBenchの3トラック全体で平均Overallスコア75.54を達成し、4つすべての設定で、比較したエンドツーエンド解析器の中で1位となった。第I段階と比べ、第II段階では、20億および40億パラメータのモデルのOverallスコアが両ベンチマークで向上した。改善は劣化を含むPureDocBenchのトラックでより大きく、Real Degradedトラックでは40億パラメータモデルで4.03ポイント向上した。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-17(UTC)
最新改訂
2026-09-17 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Document parsing converts document images into structured content and requires reliable performance across diverse layouts and acquisition conditions. Yet training corpora are biased toward common document types and clean digital pages, while expanding coverage alone does not specify how to address a parser's remaining weaknesses. We present WeVisDoc, a two-stage data-centric framework for robust end-to-end document parsing. Stage I broadens semantic, structural, and appearance coverage through heterogeneous data and structure-preserving degradation synthesis. Stage II uses a held-out probe to measure the Stage I parser's residual errors within fixed visual-structural clusters. These diagnostics guide targeted data construction and reallocation of the target-token budget. WeVisDoc-4B achieves an Overall score of 95.38 on OmniDocBench v1.6 and a mean Overall score of 75.54 across the three PureDocBench tracks, ranking first among the compared end-to-end parsers in all four settings. Compared with Stage I, Stage II improves Overall scores for the 2B and 4B models on both benchmarks, with larger gains on the degraded PureDocBench tracks, including a 4.03-point gain for the 4B model on the Real Degraded track.

arXiv ID: 2609.20423 / 要約の誤りについて