既存の汎化課題が見落とす帰納・仮説推論を調べる
What Do Current Systematic Generalization Tasks Miss? A Reasoning-Centered Analysis
この論文をやさしく読む
ひとことで言うと
学習した要素を新しく組み合わせる力を測るテストが、問題を単純にしすぎていないかを検証しています。複数の推論を同時に必要とすると、モデルの成績が大きく下がります。
何に役立つ?
モデルの汎化能力を評価する課題の設計に役立ちます。学習時より長い問題を解けるかだけでなく、必要な推論の種類も確かめるための材料です。
この研究の面白いところ
難しい課題を用意するだけでなく、単純化を1つずつ戻すと成績も戻ることを示しています。課題のどの条件が難しさを生むかを比較で調べています。
どこまで分かった?
結果は7種類のTransformerとTranSGridの4,800例に基づきます。最難関部分集合の15.8%を課題全体の成績と混同できず、人間の知能一般の測定結果でもありません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
既知の基本要素を組み合わせ直して新しい問題を解く能力である体系的汎化は、人間の知能の中心をなすが、制御された条件下で厳密に研究することは難しい。そのため既存研究では、行動のほぼ線形な合成、生産性に基づくテスト、必要な行動が明示された目標などの単純化を用いている。これらは研究を容易にする一方、この能力の本質的な側面の一部を取り除いてしまう。 単純化が何を見落とすかを明らかにするため、推論を中心とした見方を採用し、演繹、帰納、アブダクション(仮説推論)を1つの課題内で組み合わせるテスト環境TranSGridを導入する。7種類のTransformerを4,800件のTranSGrid問題で実験すると、すべてのモデルが、取り分けたテスト集合よりTranSGridで大きく低い成績となった。最大のモデルはテスト集合の79.6%を解いたが、TranSGridでは55.3%、最難関の部分集合では15.8%にとどまった。この差は訓練時の長さの範囲内でも残り、生産性だけでは体系的汎化の評価に不十分であることを示す。 さらに、残る2つの単純化をTranSGridに再導入した。一方の変種では行動の合成をほぼ線形にして帰納的な要求を減らし、もう一方では目標に行動を明示して仮説推論の要求を減らした。どちらも解答成功率はおおむねテスト集合の水準に戻り、いずれか1つの単純化だけでも、TranSGridが通常の取り分けたテスト集合と同程度の課題になることが示された。総合すると、既存課題は帰納と仮説推論の要求の一方または両方を減らしており、体系的汎化を包括的に測るには、3種類すべての推論を必要とする課題が求められる。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-16(UTC)
- 最新改訂
- 2026-09-16 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-16 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Systematic generalization, the ability to solve novel problems by recombining known atomic elements, is central to human intelligence but difficult to study rigorously under controlled settings. Existing studies therefore rely on simplifications such as approximately linear action composition, productivity-based tests, and action-explicit goals, which make systematic generalization easier to study but omit some essential aspects of this capability. To characterize what these simplifications miss, we adopt a reasoning-centered lens and introduce TranSGrid, a testbed that brings deductive, inductive, and abductive reasoning together within a unified task. Experiments with seven Transformers on 4,800 TranSGrid instances show that all models perform much worse on TranSGrid than on a held-out test set: the largest model solves 79.6% of the test set, but only 55.3% of TranSGrid and 15.8% of the hardest subset. The gap remains within the training length range, showing that productivity alone is not sufficient to evaluate systematic generalization. Additionally, we reintroduce the other two simplifications into TranSGrid: one variant makes actions compose almost linearly (reducing the inductive demand), the other makes goals action-explicit (reducing the abductive one). In both, solve rates return to roughly the test set level, showing that either simplification alone is enough to reduce TranSGrid to an ordinary held-out test set. Together, our results show that existing tasks reduce either or both of the inductive and abductive demands, and that comprehensively measuring systematic generalization requires a task that involves all three forms of reasoning.
著者のコメント
Preprint. Under review
arXiv ID: 2609.19212 / 要約の誤りについて