arXiv論文メモ
新着一覧
cs.CL · 査読状況未確認

固定観念に近い入力だけ補正して言語モデルの偏りを減らす

Acmite: Mitigating Gender Bias in LLMs through Concept-Guided Mutual Information

Tian Lan, Xiaoqing Cheng, Han Zhang, Jiang Li

この論文をやさしく読む

ひとことで言うと

固定観念に関連する概念を使って言語モデルの補正を学び、関連する入力のときだけ小さな追加モデルを動かす方法です。

何に役立つ?

ジェンダーバイアスの軽減と、それ以外の課題での能力保持を両立させるための設計として役立ちます。

この研究の面白いところ

特定の言葉の置き換えだけでなく、異なる言い回しにまたがる概念を扱います。補正を常時適用せず、入力との類似性で切り替える点も特徴です。

どこまで分かった?

報告された評価は3種類のLLMと列挙されたベンチマークです。あらゆる社会的偏りを除去したという結果ではなく、要旨には改善幅や切替判断の誤り率の具体値はありません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

大規模言語モデル(LLM)は学習データの社会的な固定観念を再現することがあり、モデルの偏りを減らす研究が幅広く行われている。しかし既存手法は、明示的に偏った例や、あらかじめ定めた集団名の置換に依存することが多く、言い回しに敏感で、多様な文脈に共通する固定観念の概念を捉えにくい。さらに重要なのは、通常、モデル出力と背後の固定観念の概念との統計的依存性を明示的にモデル化せず、偏った出力を抑えていることである。 本研究では、対象を絞って選択的に偏りを減らす、概念に導かれる軽量な枠組みAcmiteを提案する。Acmiteは固定観念を構造化した意味概念として表し、最大限界関連性(MMR)を使って補正対象の多様な概念を選ぶ。相互情報量の最小化から着想し、タスクの意味を保ちながら、トークン単位のKLダイバージェンスでこの依存性を近似する。基盤モデルを凍結したまま軽量なLoRAアダプターを学習し、推論時には入力が固定観念関連の概念と十分に似ているときだけ有効化する。それ以外では元のモデルをそのまま使う。 BBQ、CrowS-Pairs、StereoSetでAcmiteを評価し、ARC-Challenge、GSM8K、PIQAで一般的能力の保持を調べる。3種類のLLMにわたる実験は、Acmiteが相補的な評価形式でジェンダーバイアスを効果的に軽減しつつ、偏りと関係しない課題で競争力のある性能を保つことを示す。匿名化したコードとデータは https://anonymous.4open.science/r/Acmite-18E2/ で公開している。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-10-01(UTC)
最新改訂
2026-10-01 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Large language models (LLMs) can reproduce social stereotypes from their training data, motivating extensive research on model debiasing. However, existing methods often rely on explicit biased examples or predefined group-term substitutions, making them sensitive to wording and less effective at capturing stereotype concepts shared across diverse contexts. More importantly, they typically suppress biased outputs without explicitly modeling the statistical dependence between model outputs and the underlying stereotype concepts. We propose Acmite, a lightweight concept-guided framework for targeted and selective debiasing. Acmite represents stereotypes as structured semantic concepts and uses maximal marginal relevance (MMR) to select diverse concepts for debiasing. Inspired by mutual information minimization, it approximates this dependence with token-level KL divergence while preserving task semantics. A lightweight LoRA adapter is trained with the base model frozen and activated at inference time only when the input is sufficiently similar to stereotype-related concepts; otherwise, the original model is used directly. We evaluate Acmite on BBQ, CrowS-Pairs, and StereoSet, and assess general capability preservation on ARC-Challenge, GSM8K, and PIQA. Experiments across three LLMs show that Acmite effectively mitigates gender bias across complementary evaluation formats while maintaining competitive performance on bias-unrelated tasks. Anonymous code and data are available at https://anonymous.4open.science/r/Acmite-18E2/.

著者のコメント

15 pages, 0 figures

arXiv ID: 2610.01696 / 要約の誤りについて