図表の凡例を正しく読めるかを細かく診断する
LegendBench: A Diagnostic Benchmark for Legend Understanding with Counterfactual Interventions
この論文をやさしく読む
ひとことで言うと
図表の正答率だけでなく、凡例の色や記号を本当に正しく結び付けているかを調べる評価方法です。
何に役立つ?
図表を読むモデルの誤りが、凡例の読み取り、対応付け、その後の推論のどこにあるかを診断し、追加学習の対象を決めるために役立ちます。
この研究の面白いところ
同じ図表の凡例だけを統制して変えることで、意味を追えているのか、並び順などの近道に依存しているのかを区別します。
どこまで分かった?
要旨は未見データへの改善を報告しますが、改善量や対象モデル数は示していません。凡例に焦点を当てた評価であり、図表理解のすべてを単独で測るものではありません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
凡例は図表理解の基礎であり、信頼できる解釈には、凡例の項目を対応する視覚的な記号へ正しく結び付けることが必要である。視覚言語モデル(VLM)の図表理解への利用は増えているが、総合正答率だけでは凡例理解を十分に診断できない。表面的な近道によって正解できるうえ、凡例固有の誤りと他の推論の失敗が混同されるためである。 詳細な診断と統制された試験を可能にするため、凡例に焦点を当てたテスト事例を生成する、パラメータ化されたベンチマークと生成処理系LegendBenchを導入する。LegendBenchは二つの要素を提供する。第一に、凡例の解析、凡例と図中の記号の対応付け、凡例に条件付けた推論、凡例を考慮した回答保留にまたがる能力と課題の分類体系によって、失敗箇所を特定する。第二に、基礎となる各図表から、凡例を統制して変更した複数の変種を作る反実仮想グループ生成により、モデルが何を不変とし、何に反応するかを調べる。 LegendBenchを使って汎用VLMと図表専用モデルを評価し、能力の分布を作成した結果、凡例と記号を確実に対応付けることと、反実仮想的な変更に対する一貫性に、持続的な難点が見られた。次に、この能力分布に基づいて対象を絞った追加学習を行い、弱点に特化した介入が、特定された能力不足を効果的に埋め、未見データへも汎化することを示す。さらに反実仮想設計を利用して詳細な診断実験を実施し、情報を視覚表現へ割り当てる経路の効果、凡例の並び順を利用した近道、見え方が異なる条件での回答保留を分析する。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-21(UTC)
- 最新改訂
- 2026-09-21 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-21 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Legends are fundamental to chart understanding, as reliable interpretation requires correctly binding legend entries to corresponding visual marks. While vision-language models (VLMs) are increasingly applied to chart understanding, their legend understanding is poorly diagnosed by aggregate accuracy, which can be satisfied by superficial shortcuts and confound legend-specific errors with other reasoning failures. To enable fine-grained diagnosis and controlled testing, we introduce LegendBench, a parametric benchmark and generation pipeline that produces targeted legend-centric test cases. LegendBench contributes (1) a capability-task taxonomy spanning legend parsing, legend grounding, legend-conditioned reasoning, and legend-aware abstention to localize failures, and (2) counterfactual group generation, where each base chart yields multiple variants under controlled legend interventions to probe model invariance and sensitivity. Using LegendBench, we evaluate both general-purpose VLMs and specialized chart models and generate their capability profiles, revealing persistent bottlenecks in reliable legend-to-mark binding and counterfactual consistency. We then use these capability profiles to guide targeted fine-tuning, demonstrating that bottleneck-specific interventions can effectively close the localized capability gaps and generalize to unseen data. We further leverage our counterfactual design to conduct fine-grained diagnostic experiments, analyzing encoding-channel effects, legend-order shortcuts, and abstention under varying visibility.
arXiv ID: 2609.24172 / 要約の誤りについて