画像・音声と文章の提示順でAIの判断が変わる
Same evidence, different judgments: Evidence noncommutative in vision/speech-text conflicts
この論文をやさしく読む
ひとことで言うと
AIに矛盾する文章と画像・音声を見せると、後に置いた証拠が判断に強く効くことを調べた。
何に役立つ?
マルチモーダルAIの評価で、証拠の提示順が結果に与える影響を切り分けるのに役立つ。
この研究の面白いところ
指示と証拠の内容を変えず、二つの情報源の位置だけを交換して順序の効果を測った。
どこまで分かった?
要旨は画像・音声モデルで一貫した傾向を報告するが、モデル数や効果量は示していない。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
マルチモーダル大規模言語モデルに対し、画像や音声が添えられた文章と矛盾する場合、測定される文章への依存度には、モダリティーの好みと証拠の提示位置が混ざり得る。文章への偏りを調べた先行研究では、証拠の順番を固定したり、課題の指示まで証拠と一緒に動かしたりしていたため、順序の寄与が不明確だった。本研究は、指示と証拠の内容を固定し、二つの情報源の位置だけを入れ替える対比較で、その影響を定量化する。 画像モデルと音声モデルの双方で、矛盾する文章の後に画像や録音を置くと、回答が一貫してその視覚・音声内容へ寄った。また先行研究を再検討し、その実験設定が誤解を招く結論につながり得る理由を分析した。同じ証拠でも順番が変われば異なる判断に至り、知覚的な証拠を後に置くとその内容への依存が強くなる。この現象を著者らは、モダリティーをまたぐ証拠の非可換性と呼ぶ。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-22(UTC)
- 最新改訂
- 2026-09-22 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-22 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
For multimodal large language models, when images or speech conflict with accompanying text, measured text reliance can entangle modality preference with evidence position. Earlier studies of text bias often used a fixed evidence order or moved task instructions with the evidence, leaving the contribution of order unclear. In this paper, we use a paired comparison that keeps the instructions and evidence content fixed and swaps only the positions of the two sources to quantify this potential influence. Across vision and speech models, placing an image or recording after conflicting text consistently shifts answers toward its content. We also revisit previous studies and analyze why their experimental settings can lead to misleading conclusions. These findings reveal cross-modal evidence noncommutativity: the same evidence can lead to different judgments when its order changes, and placing perceptual evidence later can increase the model's reliance on its content.
arXiv ID: 2609.26986 / 要約の誤りについて