疎なオートエンコーダで文章分類モデルの判断を調べる
Exploring Text Classification Models with Sparse Autoencoders
この論文をやさしく読む
ひとことで言うと
文章を分類するモデルの内部から概念に対応する特徴を取り出し、予測や誤分類とのつながりを探るツールです。
何に役立つ?
考えられる用途は、分類モデルがどの特徴に反応しているかを調べ、誤りの原因の候補を探索することです。
この研究の面白いところ
特徴を取り出すだけでなく、モデルの予測や失敗と結び付けて調べる操作をSAEfarerというツールにまとめています。
どこまで分かった?
要旨にある利用者評価は博士課程学生5人による予備評価です。評価の具体的な結果や、特徴が予測の原因であることを確認する実験は記載されていません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
言語モデル(LM)の重要性が高まるにつれ、その内部の挙動をよりよく理解するため、透明性を高めることへの関心が生まれている。最近の解釈可能性研究では、疎なオートエンコーダ(SAE)を使い、LMのある層でのニューロンの活性化を、人間に理解できる特徴へ分解することに注目が集まっている。各特徴は、モデルが学んだ概念を表す。 本論文では、SAEを使って文章分類LMの挙動を分析する取り組みを紹介する。SAEの特徴とモデルの予測・誤りとの関係を探索する手法を提示し、それらを文章分類LMが学んだ概念を分析するツールSAEfarerへ統合する。博士課程の学生5人による専門家の予備評価で、SAEfarerを評価する。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-17(UTC)
- 最新改訂
- 2026-09-17 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-17 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
As language models (LMs) rise in prominence, there is interest in making them more transparent in order to better understand their internal behavior. Recent interpretability work has focused on using sparse autoencoders (SAEs) to break down neuron activations at a given layer in the LM into human-understandable features, where each feature represents a concept that the model has learned. In this paper, we share work on using SAEs to analyze the behavior of text classification LMs. We present techniques for exploring the relationships between the SAE's features and the model's predictions and errors. We integrate these techniques into SAEfarer, a tool for analyzing concepts learned by text classification LMs. We assess SAEfarer in an expert pilot evaluation with five Ph.D. students.
著者のコメント
11 pages, 13 figures. Accepted as a short paper at IEEE VIS 2026
arXiv ID: 2609.21142 / 要約の誤りについて