文字入り画像の配置計画と拡散描画を共同学習
Planning and Rendering in Concert: DeepFusion of Autoregressive Layouts and Diffusion for Visual Text Generation
この論文をやさしく読む
ひとことで言うと
画像に入れる文字の内容・位置を決める計画器と、画像を描く拡散モデルを一緒に学習する研究。
何に役立つ?
文章から文字入り画像を作る際、文字の正確さと画像内での位置の整合性を改善する方法の検討に役立つ。
この研究の面白いところ
描画の学習信号を計画器の内部表現にも伝え、推論時には領域と対応する座標・内容の結び付きを強める。
どこまで分かった?
要旨にある数値はCVTG-2KとLongText-Benchでの結果である。他の画像生成課題や利用環境での性能は分からない。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
文章指示から文字の多い画像を生成するには、文字を正しく表現すると同時に、周囲の画像と整合させる必要がある。明示的な配置計画は、どの文字をどこに置くかの構造的な手がかりになるが、適切な計画だけでは描画が忠実に実現されるとは限らない。既存の配置に基づく自己回帰・拡散モデルは、通常、計画と描画を別々に最適化しており、計画器の表現を画像生成と共同で調整できない。本研究は、自己回帰的な計画と連続的な拡散描画を共同学習する DeepFusion に基づく、自律的な文字入り画像生成器 DuetGen を提案する。 DeepFusion では、計画器のプロンプトおよび境界枠・内容の隠れ状態を条件として拡散Transformerを動かし、描画からの学習信号で、文字の計画と画像出力をつなぐ表現を調整する。共同目的関数は、自己回帰的な計画の教師信号、文字領域に重みを置いた拡散学習、座標の補助的な教師信号を組み合わせ、構造的な計画の維持、文字領域の重視、計画器表現の空間精度の改善を狙う。推論時には、段階を考慮した注意変調により、画像領域と対応する座標・内容状態との結び付きを強め、生成した計画を領域ごとに実行しやすくする。 20億パラメータの計画器と40億パラメータの単一ストリームDiTを用いたDuetGenは、CVTG-2Kで単語正確度0.8293、LongText-Benchで正確度0.938を達成し、両ベンチマークで、はるかに大きいQwen-Imageに近い成績だった。これらの結果は、共同学習した計画表現と領域別の描画が、自律的な文字入り画像生成に役立つことを示す。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-19(UTC)
- 最新改訂
- 2026-09-19 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-19 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Generating text-rich images from prompts requires both textual fidelity and the coherent integration of text into the surrounding image. An explicit layout can provide structured guidance about what text should appear and where, but a well-formed plan alone does not guarantee that the renderer will realize it faithfully. Existing layout-based AR-diffusion systems typically optimize planning and rendering separately, preventing the planner's representations from being adapted jointly with image synthesis. We introduce DuetGen, an autonomous visual text generator built on DeepFusion, which jointly learns autoregressive planning and continuous diffusion rendering. DeepFusion conditions a diffusion transformer on the planner's prompt and bbox-content hidden states, allowing rendering supervision to shape the representations connecting textual plans with visual outputs. Its joint objective combines autoregressive plan supervision, text-region-weighted diffusion learning, and auxiliary coordinate supervision to maintain structured planning, emphasize text-bearing regions, and improve the spatial precision of planner representations. During inference, Phase-Aware Attention Modulation strengthens the correspondence between image regions and their matched coordinate and content states, facilitating region-specific execution of the generated plan. With a 2B planner and a 4B single-stream DiT, DuetGen achieves 0.8293 word accuracy on CVTG-2K and 0.938 accuracy on LongText-Bench, closely matching the substantially larger Qwen-Image on both benchmarks. These results demonstrate the value of jointly learned planning representations and region-specific rendering for autonomous visual text generation.
arXiv ID: 2609.22916 / 要約の誤りについて