勾配ベースのデータ帰属法では内容より形式が優先される
Form Over Content In Gradient-Based Data Attribution Methods
この論文をやさしく読む
ひとことで言うと
学習データの関連性を勾配の似方で測る方法が、課題の意味より答えの書式に反応しているかを調べています。
何に役立つ?
LLMの訓練データを選ぶとき、意味的な有用性と出力形式の一致を取り違えないための評価です。
この研究の面白いところ
同じ課題で形式を変える場合と、違う課題で形式をそろえる場合を分離しました。同形式では補正後のコサインが約0.4、同課題でも別形式だと約0.0と報告します。
どこまで分かった?
教師あり追加学習例の比較が対象です。勾配に基づく全手法が無価値という結論ではなく、書式と意味を独立に変えたデータで頑健性を確認すべきだとしています。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
勾配の類似度を用いるデータ帰属法は、大規模言語モデルの訓練データを分析・選択するために広く使われているが、勾配類似度が実際に何を測っているかは議論されている。タスクに関係する技能を特定すると解釈する研究がある一方、表面的な形式が主な要因だと報告する研究もある。本研究は、教師ありファインチューニング例について、タスクと回答形式を独立に変えることでこの議論を整理する。ベンチマークを異なる回答形式で表現し、タスクは共有するが形式は共有しないデータセット、または形式は共有するがタスクは共有しないデータセットを作る。 勾配の整列は回答形式に従うことが分かった。同じ回答形式を共有するベンチマーク対は強く整列し、減衰を補正したコサインは約0.4だったが、異なる回答形式クラスで表現した同じベンチマーク対は整列せず、約0.0だった。この順序は、最初期の事前学習チェックポイントから事後学習まで、またモデルの規模や系列をまたいで成立した。 さらに、指示チューニングのための勾配ベースのデータ選択法LESSが公開した選択を分析したところ、各ターゲットの選択データはターゲット自身の回答形式を過剰に代表していた。したがって、勾配ベースの帰属法はタスク意味論より形式類似性を強く追跡することを示す。このような手法と勾配の意味論的解釈は、より堅牢で信頼できるものにするため、回答形式とタスクを独立に変えたデータで検証すべきである。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-17(UTC)
- 最新改訂
- 2026-09-17 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-17 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Data attribution methods using gradient similarity are widely used to analyze and select training data for large language models, but what gradient similarity actually measures is debated. Some interpret it as identifying task-relevant skills, while other work reports that surface form is the main factor. We resolve this debate for supervised fine-tuning examples by varying task and answer format independently. Specifically, we render benchmarks in different answer formats, such that datasets can share a task without a format or a format without a task. We find that gradient alignment follows the answer format, as benchmark pairs sharing an answer format align strongly (disattenuated cosine near 0.4), while same benchmarks rendered with different answer format classes show no alignment (near 0.0). We demonstrate that this ordering holds from the earliest pretraining checkpoints through post-training, and across model scales and families. We then analyze the released selections of LESS, a gradient-based data selection method for instruction tuning, and find that each target's selections over-represent the target's own answer format. Hence, we demonstrate that gradient-based attribution methods track format similarity more than task semantics, meaning that such methods, as well as the semantic interpretation of the gradient, should be tested on data where answer format and task vary independently for greater robustness and reliability.
arXiv ID: 2609.19589 / 要約の誤りについて