arXiv論文メモ
新着一覧
cs.CV / cs.CL · 査読状況未確認

SVG画像を推論途中で生成する視覚言語モデル

Multimodal Thinking with Renderable Programs

Sunli Chen, Ding Zhong, Ziqiao Ma, Jiaxin Liu, Zeyuan Yang, Hao Zhang, Lie Lu, Joyce Chai, Chuang Gan

この論文をやさしく読む

ひとことで言うと

視覚言語モデルが推論の途中でSVG形式の図を作り、その図を使って考える仕組みです。

何に役立つ?

図を使う数学的な推論や、文章と画像の両方を扱うデジタル作業のモデル設計に役立つ可能性があります。

この研究の面白いところ

SVGを画像とテキスト命令の両方として扱うことで、生成した図の内容を追いやすくしています。

どこまで分かった?

要旨での性能評価は数学的推論ベンチマークです。改善幅やほかの視覚課題への一般化は要旨には示されていません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

現在の視覚言語モデル(VLM)は画像内容の理解と文章による推論に優れるが、その構造は画像を推論の過程に組み込む発展を制限している。テキストと画像の生成を統一しようとするオムニモーダルなモデルもあるが、公開領域の視覚課題に重点を置き、ラスター画像や潜在表現を使うために扱いやすさが不足する。本研究は、拡張可能なベクター画像形式SVGの基本図形を用いて、推論課題でテキストと画像をつなぐ枠組みSVGLMを導入する。SVGが画像の記述であると同時にテキストによる命令でもあるという二面性を利用し、一般的なVLMが推論途中に画像を生成できるようにする、より簡潔で解釈しやすい方法を得る。SVGに基づく画像編集データの大規模な選別済みデータセットと、公開VLMを微調整する方法も提供する。数学的推論ベンチマークの実験では、SVGLMはSVG生成能力と、画像を使って考える能力の両方で高い性能を示した。結果は、SVGがテキストによる思考とピクセル画像の間をつなぎ、より頑健なデジタル領域のエージェントを作るための適切な媒体であることを示唆する。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-24(UTC)
最新改訂
2026-09-24 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Current vision-language models (VLMs) excel at visual content understanding and text-based reasoning, yet their structure limits the advancement of incorporating images into the reasoning chain. Though Omnimodal models have made efforts in unifying text and image generation, they focus on visual tasks in the open-domain, lacking tractability due to rasterized or latent representations of images. We introduce SVGLM, a framework that uses scalable vector graphics (SVG) primitives to connect text and image in reasoning tasks. We exploit the duality of SVG as both image description and text instructions, yielding a more compact, interpretable solution to equip general VLMs with the capability of generating images within the reasoning process. We provide a large curated dataset of SVG-based image editing dataset, as well as the paradigm to tune open-source VLMs. Experiments on a mathematical reasoning benchmark demonstrate that SVGLM achieves strong SVG generation power as well as think-with-image intelligence. Our results highlight SVG as a suitable medium for building more robust digital domain agents, bridging the gap between text-based thinking and pixel-based images.

arXiv ID: 2609.30130 / 要約の誤りについて