画像推論の内部トークンに異なる情報を担当させる
MoLE: Mixture of Latent Experts for Complementary Visual Reasoning
この論文をやさしく読む
ひとことで言うと
画像を内部で考える複数のトークンが同じ情報ばかり拾わないよう、別々の視覚情報を担当させ、後からまとめる方法です。
何に役立つ?
潜在トークン数に制約がある視覚言語モデルで、内部計算の重複を減らし、視覚推論の性能を高める設計に役立ちます。
この研究の面白いところ
どの情報を見るかだけでなく、どう変換するかも専門家ごとに制御します。役割を人が決めず、学習中に視覚情報を潜在経路へ通すことで分担を促しています。
どこまで分かった?
結果は五つの評価課題の平均です。4.9、3.6、9.2は要旨のスコア差の表記を保持しており、百分率の改善としては扱っていません。潜在経路を消したときの低下も、評価した条件でのアブレーション結果です。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
潜在視覚推論は、視覚言語モデルに連続的な中間状態を与え、明示的な文章の推論過程や、画像への繰り返しの操作を使わずに視覚的証拠を処理できるようにする。しかし既存手法では、複数の潜在トークンが共有された値射影を通じて同じ視覚的証拠へアクセスすることが多く、相補的な視覚情報を取り出す仕組みがない。そのため、潜在トークンの予算を増やすだけでは、冗長な潜在表現が生まれ得る。 本研究では、効果的な潜在推論には、異なる潜在トークンが相補的な視覚情報を抽出し、専門化した視覚の専門家として機能することを促すべきだと論じる。この着想に基づき、各潜在視覚専門家がどの証拠を見るかと、その証拠をどう変換するかの両方を制御するMixture of Latent Expertsの枠組みMoLEを提案する。MoLEは証拠抽出の間、潜在視覚専門家を分離し、専用の潜在要約専門家によって、それぞれの相補的な表現を集約する。2段階の学習では、まず視覚的証拠をこの潜在経路に強制的に通し、その後、直接の視覚アクセスを復元する。専門家の役割を事前に定めることも、中間の視覚的な教師信号も必要としない。 五つの視覚推論ベンチマークで、MoLEは平均スコア78.6を達成し、データをそろえた教師あり微調整を4.9、同じ潜在予算で評価した最も強い潜在視覚推論の比較手法を3.6上回った。表現分析では、潜在状態間の類似度が低く、視覚的注意がより多様であることが示された。一方、潜在経路をマスクすると平均性能が9.2低下した。これらの結果は、潜在トークンの数を増やすだけよりも、潜在計算を専門化させる方が有効であることを示している。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-10-01(UTC)
- 最新改訂
- 2026-10-01 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-10-01 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Latent visual reasoning equips vision--language models with continuous intermediate states that can process visual evidence without explicit textual reasoning traces or repeated image operations. However, existing methods often allow multiple latent tokens to access the same visual evidence through shared value projections, providing no mechanism for them to extract complementary visual information; simply increasing the latent budget can therefore yield redundant latent representations. We argue that effective latent reasoning should encourage different latent tokens to extract complementary visual information, and thereby act as specialized visual experts. Based on this insight, we propose MoLE, a Mixture of Latent Experts framework that controls both what visual evidence each latent visual expert observes and how it transforms that evidence. MoLE isolates latent visual experts during evidence extraction and uses dedicated latent summary experts to aggregate the complementary representations of latent visual experts. A two-stage training pipeline first forces visual evidence through this latent pathway and then restores direct visual access, requiring neither predefined expert roles nor intermediate visual targets. Across five visual reasoning benchmarks, MoLE achieves an average score of 78.6, outperforming data-matched supervised fine-tuning by 4.9 and the strongest evaluated latent visual reasoning baseline at the same latent budget by 3.6. Representation analyses show lower latent-state similarity and more diverse visual attention, while masking the latent pathway reduces average performance by 9.2. These results demonstrate that specializing latent computation is more effective than merely increasing the number of latent tokens.
arXiv ID: 2610.01917 / 要約の誤りについて