arXiv論文メモ
新着一覧
cs.CV / cs.AI · 査読状況未確認

誤認識の連鎖を抑える一段階の文字画像高解像度化

TOLA: Text-aware One-Step Latent Adaptation for Diffusion-based Text Image Super-Resolution

Yike Xu, Yue Shi, Yong Guo, Jiezhang Cao

この論文をやさしく読む

ひとことで言うと

ぼやけた文字画像を読みやすくする際、最初の文字認識ミスが画像復元に広がるのを防ぐ方法。

何に役立つ?

文字画像の高解像度化で、処理時間と誤った文字の生成を抑えたい場合に役立つ可能性がある。

この研究の面白いところ

文字の候補を一度だけ条件として使い、信頼度の低いOCR結果を抑えた上で、欠けた筆画を別の軽量な補正部で修復する。

どこまで分かった?

性能の主張はCTR-TSR-TestとRealCE-200での評価に基づく。ほかの種類の文字画像で同じ差が出るかは要旨に記載されていない。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

文字画像の超解像は、劣化の仕方が分からない画像から、見た目が忠実で読める文字を復元することを目指す。既存の拡散モデルによる方法は、高解像度画像か文字に関する事前情報を複数段階で予測することが多く、計算量と推論の遅延が大きい。さらに重要な問題として、誤った文字の事前情報がノイズ除去の過程に繰り返し注入されると、画像予測と文字予測が互いを強め合い、初期の認識誤りが鮮明だが意味的に間違った文字へと増幅される。 この問題に対し、画像と文字の拡散を繰り返さない、文字情報を考慮した一段階の潜在表現適応法TOLAを提案する。主な構成要素は二つある。第一に、信頼度で重み付けした文字条件付け部は、意味的な条件を一度だけ作り、信頼できないOCR予測が画像の再構成を汚染する前に抑える。第二に、軽量な潜在残差補正部は、構造的な残差誤りを明示的に推定・補正し、欠けたり歪んだりした文字の線を復元する。 広範な実験では、CTR-TSR-Testの4倍拡大とRealCE-200の両ベンチマークで、すべての評価指標において最先端の性能を示した。特にCTR-TSR-Testでは、既存の拡散モデルによる文字超解像法よりPSNRが常に少なくとも2.72 dB高かった。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-24(UTC)
最新改訂
2026-09-24 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Text image super-resolution (TSR) aims to recover visually faithful and readable text under unknown degradations. Existing diffusion-based methods typically rely on multi-step prediction of either the high-resolution image or its text prior, resulting in prohibitive computational cost and inference latency. More critically, an erroneous text prior may be repeatedly injected into the denoising process, causing image and text predictions to reinforce each other and progressively amplify an early recognition error into a sharp yet semantically incorrect character. To address these limitations, we propose TOLA, a Text-aware One-step Latent Adaptation framework without iterative image-text diffusion. TOLA consists of two key modules. First, a confidence-weighted text conditioning module constructs the semantic condition only once and suppresses unreliable OCR predictions before they contaminate image reconstruction. Second, a lightweight latent residual correction module explicitly estimates and corrects the structured residual errors to recover missing or distorted stroke details. Extensive experiments demonstrate our state-of-the-art performance across all evaluation metrics on both CTR-TSR-Test ($\times 4$) and RealCE-200 benchmarks. It is worth noting that our TOLA consistently surpasses existing diffusion-based TSR methods by at least 2.72 dB in PSNR on CTR-TSR-Test.

著者のコメント

18 pages, 11 figures, including appendices

arXiv ID: 2609.29240 / 要約の誤りについて