arXiv論文メモ
新着一覧
cs.CV · 査読状況未確認

画像の小さな変化による視覚言語モデルの回答反転を抑える

Looks the Same, Answers Differently: Flip-Direction Steering for Robust Vision-Language Reasoning

Yeonsung Jung, Joonhyun Jeong, Hoang Pham, Joowon Kim, Yoonsik Park, Viet Dac Lai, Eunho Yang

この論文をやさしく読む

ひとことで言うと

画像がほぼ同じでもAIの回答が変わる問題に対し、推論時の内部状態を調整して回答の安定性を高める方法です。

何に役立つ?

考えられる用途は、撮影条件や画像処理の小さな違いに影響されにくい視覚言語モデルの評価と改善です。

この研究の面白いところ

反転した画像対から活性化の方向を推定し、不確実な生成段階だけに作用させます。元の回答の回復と、もともと安定した回答の保持を分けて測ります。

どこまで分かった?

要旨に示された結果は9通りのデータセット・変化の組み合わせを含む18設定での比較です。改善幅の数値や、すべての視覚変化への一般化は要旨に記載されていません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

視覚言語モデルは高い視覚推論性能を示すが、通常の撮影や画像処理で生じるわずかな変化でも、画像がほぼ同じに見えるのに推論の経路が変わることがある。長い文章を生成する場合、活性化の変化が復号の各段階で蓄積し、推論に使うトークンが次第に変化して、最後には回答が反転することがある。 この不安定さに対し、学習を追加せず推論時に適用するFlipDir(Flip-Direction Steering)を提案する。元画像と回答を反転させる入力との対から、反転を引き起こす活性化の低ランク部分空間を推定し、復号中の隠れ状態を選択的に調整する。マージンに基づくゲートにより、不確実な復号段階だけでその部分空間の影響を弱め、元の予測を回復しつつ、安定した予測を保持する。 固定されたテストセットでの正解率や一貫性だけを超えて頑健性を評価するため、VisFlipというベンチマークの枠組みも導入する。対象モデルと画像変化の設定ごとに評価群を構成し、元の予測の回復と安定した予測の保持を別々に評価する。VisFlipは、科学的推論、ロボットの場面理解、医療画像の質問応答における典型的な微小な視覚変化を含む、9通りのデータセットと変化の組み合わせからなる。18の設定での実験では、FlipDirは回復と保持を組み合わせた指標で既存手法を一貫して上回った。コードは今後公開予定である。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-23(UTC)
最新改訂
2026-09-23 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Vision-language models (VLMs) achieve strong visual reasoning performance, yet subtle changes from routine image capture and processing can alter their reasoning trajectories even when images appear nearly identical. In long-horizon generation, the resulting activation shifts may accumulate across decoding steps, progressively altering reasoning tokens and ultimately changing the final answer, a phenomenon referred to as answer flips. To address this instability, we propose FlipDir (Flip-Direction Steering), a training-free inference-time method that estimates a low-rank flip-inducing activation subspace from contrastive pairs of original and answer-flipping inputs and selectively steers hidden states during decoding. A margin-based gate limits subspace attenuation to uncertain decoding steps, recovering original predictions while preserving stable ones. To evaluate robustness beyond accuracy or consistency on fixed test sets, we introduce VisFlip, a benchmark framework that constructs evaluation groups for a target model and visual variation setting to separately assess recovery of original predictions and preservation of stable ones. VisFlip spans nine dataset-variation combinations across scientific reasoning, robot-scene understanding, and medical VQA, covering subtle visual variations common in each domain. Experiments across 18 settings demonstrate that FlipDir consistently outperforms existing methods on the combined recovery and preservation metric. We will make our code publicly available.

著者のコメント

27 pages

arXiv ID: 2609.28851 / 要約の誤りについて