論文の図を編集可能なスライドへ再構成するAIを評価
ReFigBench: Benchmarking Scientific Figure Reconstruction as Editable PowerPoint Artifacts
この論文をやさしく読む
ひとことで言うと
論文図をPowerPointに作り直すAIを、見た目だけでなく文字や接続線を編集できるかまで含めて評価する。
何に役立つ?
図の再利用やスライド作成を支援するAIを選ぶ際、画像の類似度だけでは分からない成果物の品質を評価するのに役立つ。
この研究の面白いところ
同じモデルでも周囲の実行環境によって専用ワークフローの効果が逆転し、見た目の好評価と文書構造の保持も両立しない場合を示している。
どこまで分かった?
対象は科学的概要図1,000点と十構成である。専用ワークフローでは全構成でネイティブコネクタが失われており、見た目の改善だけで編集可能性が改善したとは言えない。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
マルチモーダルなコーディングエージェントには、視覚的な入力から利用可能な成果物を作ることが期待されている。エージェントは、モデルを取り巻くツール、文脈管理、実行環境の層であるハーネスを介して動作する。既存の評価では、短いツール呼び出し、APIの実行記録、スクリーンショットの見た目の類似性などを切り離して測ることが多い。しかし、こうした代理指標で低い得点を得ても、モデルの視覚認識が悪かったのか、計画が悪かったのか、ハーネスに妨げられたのかは分からない。 本研究は、元画像を、文字、接続関係、配置、文書本来の構造を保った編集可能なPowerPointスライドへ変換する、科学的な概要図の再構成を扱う。arXiv論文から取得し、出典を完全に追跡できる実際の概要図1,000点を用いたベンチマークと評価の枠組みReFigBenchを導入する。四つのモデル系統のコーディングエージェントが、直接のコード生成とPPTX専用ワークフローという二つの方式で各図を再構成する。最も強力なモデルは二つの商用ハーネスでも実行し、計十通りの構成を評価する。評価には、決定的な成果物チェック、二つのモデル系統の評価者による反復自動採点、条件を伏せた人による比較を組み合わせる。 視覚認識は依然としてボトルネックであり、繰り返し描画して確認しても、その改善は部分的にとどまる。ワークフローへの労力が品質につながるかどうかは、モデルとハーネスの組合せに依存する。同じモデルでも、専用ワークフローにより一方のハーネスでは改善し、もう一方では悪化する。また、直接生成のプロンプトが同一でも、ハーネスによって得点が変わる。専用ワークフローではすべての構成で文書本来のコネクタが失われるが、人の評価者はそれでも大半の対戦比較でその描画を好む。最も強力なエージェントでさえ評価基準の上限には届かない。これらの結果は、実用的なマルチモーダル文書エージェントの中心課題が、忠実な見た目と編集可能性の緊張関係にあることを示す。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-16(UTC)
- 最新改訂
- 2026-09-16 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-16 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Multimodal coding agents are expected to turn visual inputs into usable artifacts, and they act through a harness, the layer of tools, context management, and execution environment around the model. Existing evaluations often isolate short tool calls, API traces, or screenshot resemblance, and a low score under these proxies cannot say whether the model saw poorly, planned poorly, or was failed by its harness. We study scientific overview figure reconstruction, an agent task in which a source image must become an editable PowerPoint slide that preserves text, topology, layout, and native document structure. We introduce ReFigBench, a benchmark and evaluation framework built on 1,000 real overview figures retrieved from arXiv papers with full provenance. Coding agents from four model families reconstruct every figure under two workflows, direct code generation and a specialized PPTX workflow, and the strongest model runs inside two commercial harnesses, yielding ten configurations. Evaluation combines deterministic artifact checks, repeated automated scoring by judges from two model families, and blinded human comparisons. Perception remains a bottleneck that iterative rendering only partly repays. Whether workflow effort converts into quality depends on the model together with its harness, since the same model gains from the specialized workflow inside one harness and loses inside the other, and the harness shifts scores even under an identical direct prompt. The specialized workflow erases native connectors in every configuration, human judges still prefer its renderings in most matchups, and even the strongest agent falls short of the rubric ceiling. These results expose the tension between fidelity and editability as the central challenge for practical multimodal document agents.
著者のコメント
31 pages, 7 figures, including appendices
arXiv ID: 2609.18844 / 要約の誤りについて