抽象推論を五つの能力で測るPotARCin
PotARCin: Multi-Dimensional Evaluation of Skill Acquisition in Abstract Reasoning Tasks
この論文をやさしく読む
ひとことで言うと
ARC課題の正解格子だけでなく、規則の説明や逆変換など五つの能力を測る評価基準。
何に役立つ?
AIモデルが規則を本当に扱えるかを、単一の正解率より詳しく調べる用途に役立つ。
この研究の面白いところ
標準評価との間に25~52ポイントの差があり、モデルが述べた規則と自分の回答が矛盾する例も見つかった。
どこまで分かった?
評価は五つのモデルと記載されたARC由来の集合、手作業のP-ARCに基づく。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
Abstraction and Reasoning Corpus(ARC)は、AIモデルの一般的な抽象推論と流動性知能を評価する有力な基準になっている。しかし標準的なARCの評価は、試験入力に対して正しい出力の格子を作るという一つの能力だけを見る。著者らは、この狭い形式では、本当の抽象的な技能の獲得によって可能になる多様な能力を評価できないと主張する。そこでPotARCinを導入し、課題の基礎にある抽象規則の理解を、定義、分類、制約付き生成、編集、逆変換の五つの側面で評価する。PotARCinはプログラムによって新しい課題事例を生成し、既存のARC課題の入力を変換するため、固定された入出力の組を超える動的な生成標本を作れる。ARC-AGI-1の学習用集合で五つの最先端モデルを評価すると、標準的なARC評価とPotARCin評価の間に25~52ポイントの性能差が見られ、多面的な評価では標準的な正解率で同順位のモデルの順位も変わった。さらに、生成標本、壊し方の難度、自己一貫性の影響を調べると、モデルは規則を正しく述べた場合でも、自分で形式化した規則と矛盾することがしばしば分かった。また、手作業で作成した未使用の試験集合P-ARCも導入した。この集合では、五つの側面すべてを通じた正解率は1~8%であり、抽象推論能力をより総合的に評価する重要性を示す。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-23(UTC)
- 最新改訂
- 2026-09-23 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-23 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
The Abstraction and Reasoning Corpus (ARC) has become a prominent benchmark for evaluating general abstract reasoning and fluid intelligence in AI models. Yet standard ARC evaluation considers only a single capability: producing the correct output grid for a test input. We argue that this narrow format fails to evaluate the diversity of abilities that genuine abstract skill acquisition should enable. We introduce PotARCin, a benchmark that extends ARC by assessing understanding of a task's underlying abstract rule across five dimensions: Definition, Classification, Constrained Generation, Editing, and Inversion. PotARCin employs programmatic methods to generate new task instances and transform given inputs for a given ARC task, enabling dynamic generative sampling beyond fixed input-output pairs. Across five state-of-the-art models evaluated on the ARC-AGI-1 training set, we observe a 25-52 percentage-point performance gap between standard ARC evaluation and evaluation on PotARCin, and find that multi-dimensional evaluation reorders models that standard accuracy ranks alike. We further investigate effects of generative sampling, difficulty of corruption types, and questions of self-consistency, showing that models frequently contradict their own formalized rule even where they have stated it correctly. We also introduce P-ARC, a held-out hand-crafted test set, on which models achieve 1-8% accuracy across all five dimensions, underscoring the importance of more holistic evaluations of abstract reasoning capabilities.
arXiv ID: 2609.27288 / 要約の誤りについて