arXiv論文メモ
新着一覧
cs.CV · 査読状況未確認

例示学習が視覚言語モデルの性別偏りに与える影響

Gender Bias in Vision-Language In-Context Learning

Tong Xiang, Noa Garcia, Yuta Nakashima

この論文をやさしく読む

ひとことで言うと

文脈内で見せる画像の性別が、視覚言語モデルの出力の偏りを動かすと示した。

何に役立つ?

考えられる用途は、画像説明などの例示データを選ぶ際の偏り評価である。

この研究の面白いところ

通常の品質指標では見えない偏りを測り、実画像を合成画像に替える介入も試す。

どこまで分かった?

六モデル、三課題、四データセットでの結果で、視覚質問応答では同じ効果は見られなかった。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

文脈内学習(ICL)によって大規模視覚・言語モデル(LVLM)は例のパターンを使って課題を解けるが、社会的な偏りを増幅する可能性は十分に調べられていない。VL-BICLEという評価枠組みを用い、六つのICL設定、三つの課題、四つのデータセットで、ICLが性別の偏りにどう影響するかを系統的に調べる。六つのLVLMでの実験では、性別を示す例がモデルの偏りをその性別へ動かし、反対の性別での性能を不釣り合いに低下させることが分かった。この効果は画像説明と代名詞の予測で見られたが、視覚質問応答では見られず、出力が性別を表す言語を含む課題で偏りが変わることを示す。類似度による例の検索は学習用の例群にある性別の不均衡を引き継ぎ、偏りを減らす利点はなかった。また通常の品質指標は、この偏りの変化を捉えられなかった。対策として、文章の説明を変えずに、文脈内の実画像をStable Diffusionモデルによる合成画像に替えた。この単純な介入は、画像説明の品質を下げずに性別の偏りを減らした。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-23(UTC)
最新改訂
2026-09-23 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

In-context learning (ICL) enables large vision-language models (LVLMs) to perform tasks by following patterns from in-context examples, yet its potential to amplify societal biases remains underexplored. We systematically investigate how ICL influences gender bias in LVLMs through VL-BICLE, an evaluation framework comprising six ICL settings, three tasks, and four datasets. Our experiments on six LVLMs reveal that gendered ICL demonstrations act as a directional force, shifting model bias toward the demonstrated gender through a cross-gender mechanism that disproportionately degrades performance on the opposite gender. This effect appears in image captioning and pronoun prediction but not in visual question answering, indicating that gendered ICL influences bias only when the task output involves gendered language. Similarity-based retrieval methods inherit the training pool's gender imbalance and offer no debiasing advantage, while standard quality metrics remain blind to these bias shifts. To mitigate this bias, we replace real in-context images with synthetic ones from stable diffusion models while keeping captions unchanged. This simple intervention reduces gender bias without degrading caption quality.

著者のコメント

Accepted to ECCV 2026

arXiv ID: 2609.27682 / 要約の誤りについて