arXiv論文メモ
新着一覧
cs.CL / cs.AI · 査読状況未確認

問題が難しくなる理由を自然言語の仮説で説明する

What Makes Something Hard(er)? Explaining Question Difficulty in Natural Language

Peng Cui, Qiaoyuan Zheng, Rudolf Debelak, Mrinmaya Sachan

この論文をやさしく読む

ひとことで言うと

難易度の点数だけでなく、問題文のどんな特徴が難しさを生むかを文章で説明し、別の問題や編集実験で確かめる方法です。

何に役立つ?

モデルの能力差を測る評価問題を設計したり、特定の要因を変えて難易度を調整したりする際に役立つ可能性があります。

この研究の面白いところ

仮説を読んで納得するだけで終わらず、未知の問題の難易度予測と、仮説に沿った問題編集の両方で検証しています。

どこまで分かった?

ここでの難易度は多数のLLMの回答から推定したもので、人間の難易度を直接測ったものではありません。編集実験が支える因果的な解釈も、検証した仮説と問題の範囲に基づきます。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

難易度は問題の最も基本的な性質の一つであり、能力の異なるモデルをその問題が意味のある形で区別できるかを決める。現在、さまざまな方法で難易度を自動的に推定・予測できるが、得られるのは記述的な数値一つだけであり、そもそも何が問題を難しくしているのかという根底の要因は説明されない。 本研究では、ある問題が別の問題より難しい理由を説明する自然言語の仮説を、自動生成して検証するデータ駆動型の方法を提案する。まず、多数のLLMの回答を用い、項目反応理論によって各問題の難易度を推定する。次に、対照的な易しい問題群と難しい問題群を抽出し、LLMにその違いの説明候補を提案させ、取り置いた問題で検証して選択する。 数学的推論、論理的推論、常識推論にまたがる三つのデータセットでの実験から、この方法は解釈可能で予測力のある仮説を生成することが示された。仮説単独でも、未知の問題の難易度を高度なブラックボックス型の難易度回帰モデルに匹敵する、またはそれ以上の性能で予測できる。追加特徴量として使えば、その回帰モデルをさらに改善でき、既存モデルが捉えられない難易度の手掛かりを発見していることが示唆される。さらに、仮説に従って問題を編集すると、測定された難易度が期待する方向へ変化することを実証した。これは、発見された仮説が事後的な記述ではなく、因果的に妥当な難易度要因であることを示す。本手法はこのように、純粋に記述的だった難易度スコアを、行動に移せる記述へ変換する。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-10-01(UTC)
最新改訂
2026-10-01 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Difficulty is one of the most fundamental properties of a question: it determines whether the question can meaningfully discriminate between models of differing ability. Although a variety of methods can now estimate or predict difficulty automatically, they yield only a single descriptive number, with no account of the underlying factors that make a question difficult in the first place. In this work, we propose a data-driven approach that automatically generates and validates natural-language hypotheses explaining what makes one question harder than another. We first estimate each item's difficulty from the responses of a large pool of LLMs using Item Response Theory. We then sample contrasting sets of easy and hard questions and prompt an LLM to propose candidate explanations of the difference, which are subsequently validated and selected on held-out questions. Experimental results across three datasets spanning mathematical, logical, and commonsense reasoning show that our method produces interpretable and predictive hypotheses. On their own, they predict the difficulty of unseen questions competitively with, or better than, advanced black-box difficulty regressors; used as additional features, they further improve those regressors, implying that they discover difficulty signals that existing models fail to capture. Moreover, we demonstrate that editing questions according to a hypothesis can shift their measured difficulty in the expected direction, indicating that the discovered hypotheses are causally valid difficulty factors rather than post-hoc descriptions. Our approach thus turns a purely descriptive difficulty score into actionable statements.

arXiv ID: 2610.01627 / 要約の誤りについて