arXiv論文メモ
新着一覧
cs.CL / cs.AI · 査読状況未確認

言語モデルの価値の好みを重みの方向として分ける

Geometry of Values: Task Vector Composition for Ethical Preference Alignment in Language Models

Utkarsh Agarwal, Monojit Choudhury

この論文をやさしく読む

ひとことで言うと

AIが誠実さ、正義、自律のどれを優先するかを二択問題で調べ、その好みを学習による重みの変化から切り分ける研究です。

何に役立つ?

考えられる用途は、言語や選択肢の順番による偏りを評価し、指示への従い方と価値の優先順位を分けて分析することです。

この研究の面白いところ

一般的な指示追従の変化を差し引くようにベクトルを直交化し、残った価値選好の方向を使って反対の立場へ変えています。

どこまで分かった?

98%超はこのデータセット上の正解率であり、道徳的に正しい判断の普遍的な割合ではありません。評価は特定の三組の価値対立とモデルに基づきます。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

大規模言語モデル(LLM)は、対立する道徳的価値を比較衡量しなければならない用途にますます導入されているが、高性能なモデルにも、隠れた偏りや言語をまたぐ指示追従の脆さがある。本研究では、誠実さ対正義、正義対自律、そして自律対誠実さという三つの価値の組み合わせを対象とする、二択のジレンマ12000例のデータセットを導入する。さらにヒンディー語、アラビア語、スペイン語、中国語への翻訳を用意し、言語間の振る舞いを調べる。 GPT-5-miniのベンチマークでは、方針を与えないと、5言語すべてで自律より誠実さを一貫して優先することが分かった。Llama-3.2の1B・3Bモデルは、最初の選択肢を選ぶ強い偏りを示した。しかし、通常のファインチューニングと直接選好最適化(DPO)によるファインチューニングのいずれも、この偏りを効果的に除去し、正解率を98%超に高めた。 データセット内の相関を学ぶ効果を抽象的な価値から分離するため、タスクベクトルの転移に基づく実験を提案する。ある価値を優先する方向のタスクベクトルを計算した後、一般的な指示追従のベクトルに対して直交化する。実験では、この方法が特定の価値選好の方向を切り分けるのに有効であり、それを用いたタスク算術によって、逆の立場を取るモデルを得られることを示す。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-17(UTC)
最新改訂
2026-09-17 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Large Language Models (LLMs) are increasingly deployed in applications that must weigh clashing moral values, yet even strong models exhibit hidden biases and brittle instruction-following across languages. We introduce a 12,000-instance dataset of two-option dilemmas covering pairwise three value conflicts: Honesty vs. Justice, Justice vs. Autonomy, and Autonomy vs. Honesty, along with their translations into Hindi, Arabic, Spanish, and Chinese, to probe cross-lingual behavior. Benchmarking on GPT-5-mini reveals that it consistently favors Honesty over Autonomy across all five languages when no policy is given. The Llama-3.2-1/3B models exhibit strong first-option bias; however, both plain fine-tuning and Direct Preference Optimization fine-tuning effectively remove this bias, increasing accuracy to greater than 98%. In order to decouple the effect of learning correlations in the dataset from abstract values, we propose a task vector transfer based experiment where after computing the task vectors for a direction of value preference we orthogonalize it with respect to the general instruction following vector. Our experiment shows that this method is effective in isolating the direction of the specific value preference that can successfully be used to conduct task arithmetic to obtain a model with the opposite stance.

著者のコメント

Accepted at the Pluralistic Alignment Workshop @ ICML 2026, Seoul, South Korea. https://icml.cc/virtual/2026/75692

arXiv ID: 2609.21094 / 要約の誤りについて