arXiv論文メモ
新着一覧
cs.AI / cs.LG · 査読状況未確認

強化学習後の言語モデルの合成的推論

Compositional Reasoning in Language Models under Reinforcement Learning Post-Training

Yu He and Yingxi Li and Yifei Wang and Ellen Vitercik

この論文をやさしく読む

ひとことで言うと

言語モデルに個別の技能を教えても、その組み合わせ問題を解けるとは限らないことを調べています。逆に、組み合わせ問題で学ぶと個別問題へは移りやすい傾向がありました。

何に役立つ?

複雑な課題を解くモデルの訓練データを設計し、技能の転用を評価する材料になります。個別技能の正解率だけでは組み合わせ能力を判断しにくいことを示します。

この研究の面白いところ

依存グラフで合成の複雑さを三段階に整理し、報酬が明確なデータ構造課題で非対称性を検証しています。理論的な説明と実験を組み合わせています。

どこまで分かった?

主な実証はデータ構造課題です。実用的なツール呼び出しへの広がりは予備研究の段階で、すべての実務課題に同じ傾向があると確定したわけではありません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

合成的推論は現実の問題解決に不可欠です。学習データには限りがあるため、モデルは学習済みの技能を新しい組み合わせで構成して一般化する必要があります。強化学習(RL)などの事後学習手法は言語モデル(LM)の推論能力を大きく高めてきましたが、合成的推論にどのような影響を与えるかは十分に理解されていません。 本研究では、合成的推論を形式化する依存グラフの枠組みを提案し、複雑さが増す三つの合成性の水準を導きます。実証評価では、報酬を決定論的に計算でき、構成が明確なデータ構造タスクをこの枠組みに当てはめました。その結果、分解された技能から合成された課題への移行は安定して起こらない一方、合成課題での学習は分解課題へ比較的移行しやすいという、分解から合成への非対称性が一貫して見られました。 この非対称性について理論的な説明を与え、さらに、系列長の外挿、構造分布の変化、未知の技能を必要とする課題への転移という条件で、合成的一般化を評価しました。最後に、現実的なツール呼び出しベンチマークで予備的な調査を行い、分解から合成への非対称性が実用的な設定にも及ぶ可能性を示す初期的な証拠を得ました。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-16(UTC)
最新改訂
2026-09-16 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Compositional reasoning is critical for real-world problem solving: since training data is necessarily limited, models must generalize by composing learned skills in new ways. While post-training methods such as reinforcement learning (RL) have substantially improved the reasoning abilities of language models (LMs), their effects on compositional reasoning remain less well understood. We propose a dependency-graph framework to formalize compositional reasoning, yielding three levels of compositionality with increasing complexity. Empirically, we instantiate this framework with data-structure tasks, which provide deterministic reward computation and clear compositional structure. We find a consistent decomposed-to-composed asymmetry: decomposed-skill training does not reliably transfer to composed tasks, whereas composed-task training transfers more readily back to decomposed tasks. We provide theoretical explanation for this asymmetry, and further evaluate compositional generalization under length extrapolation, structural distribution shift, and transfer to tasks requiring unseen skills. Finally, we present a pilot study on real-world tool-calling benchmarks, showing preliminary evidence that the decomposed-to-composed asymmetry can extend to practical settings.

arXiv ID: 2609.19465 / 要約の誤りについて