arXiv論文メモ
新着一覧
cs.AI · 査読状況未確認

金融文書の自動承認を支える信頼度の較正

Perception, Layout, and Validation: Calibrated Confidence for Reliable Straight-Through Processing of Financial Documents

Yichao Jin, Yushuo Wang, Yuxuan Han, Kwan Ching Yee Sonia, Weiyang Song, Chiu Jin-Chun Kent, Wong Chong Hwee, Wong Tiong Kiat, Kenneth Zhu Ke, Jingyuan Zhao

この論文をやさしく読む

ひとことで言うと

財務書類から読み取った値を自動承認するため、画像認識、配置、検証の三つから信頼度を組み立てます。

何に役立つ?

人手確認へ回す項目と自動処理する項目を、抽出モデル自身の自己申告より適切に分けるための方法です。

この研究の面白いところ

三つの説明可能な信号に適合リスク制御を組み合わせます。三データセット・二モデル系で、正誤の分離と指定リスク下の自動承認率を比較します。

どこまで分かった?

49〜72%の自動承認は、目標誤り率10%未満という評価条件での結果です。誤りゼロの保証ではなく、実務上許容できる誤り率を自動的に満たすと判断することはできません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

金融文書から抽出したキーと値の項目を人の確認なしで一貫して自動処理するストレートスルー処理(STP)には、較正された確率と、自動承認される項目群の残留誤りに対する上限保証が必要である。近年の視覚言語モデル(VLM)は、そのままキーと値を抽出できる一方、言葉で表明する信頼度は当てにならず、項目の正しさとの対応も弱い。 本論文では、知覚、レイアウト、検証という解釈可能な3つの経路に分解した信頼度層を導入する。このスコアを最終段階のコンフォーマル・リスク制御と組み合わせることで、金融文書の信頼できるSTPに利用できる。実際の請求書、合成請求書、広告購入フォームを含む3つの公開データセットで、異なる2系列のVLM、Qwen3.6-27BとGemini-3.1-Flash-Liteを用いて手法を検証した。 提案する分解スコアは、正しい抽出と誤った抽出の分離を一貫して改善する。設計した3経路すべてが寄与し、AUROCは、VLMが言語化した信頼度による0.54〜0.74から、0.90〜0.99へ大幅に上昇する。産業導入にとって特に重要なのは、これによって利用に耐えるSTPが可能になることである。目標誤り率を10%未満に設定したリスク制御の下では、VLM本来の信頼度で承認できたのは項目の0.1〜7.0%にとどまった。これに対して提案法は、承認対象群の経験的誤り率を目標以下に保ちながら、49〜72%の項目を自動承認する。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-17(UTC)
最新改訂
2026-09-17 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Straight-through processing (STP) on extracted key-value fields from financial documents without human review requires a calibrated probability together with a bounded guarantee on the residual error of the auto-approved tier. The emergence of modern Vision Language Models (VLMs) provides an out-of-the-box capability for extracting the key-values, but their verbalized confidence signals are unreliable and weakly track field correctness. This paper introduces a decomposed confidence layer along three interpretable channels, including perception, layout, and validation. Together with a final conformal risk control, the score can be used for reliable STP of financial documents. The method is validated on three public datasets covering real invoices, synthetic invoices, and ad-buy forms, using two different VLM families (Qwen3.6-27B and Gemini-3.1-Flash-Lite). Our decomposed score consistently improves the separation of correct from incorrect extractions, substantially raising the AUROC from 0.54-0.74 for VLM verbalized signals to 0.90-0.99 with contributions from all three designed channels. Crucially for industrial deployment, this enables usable STP. The native VLM confidence signals could clear only 0.1%-7.0% of fields under risk control at a target error of <10%. In contrast, the proposed method auto-approves 49-72% of fields while holding the empirical error of the accepted tier at or below the target.

arXiv ID: 2609.20110 / 要約の誤りについて