arXiv論文メモ
新着一覧
stat.ML / cs.LG · 査読状況未確認

言語モデルの複雑な問題解決を統計的な制御として整理

Complex Problem Solving in Large Language Models: A Statistical Control Survey and Diagnostic Framework

Jiazhang Cai, Tao Wang, Ruidong Zhang, Siyuan Li, Terry Ma, Luyang Fang, Haoran Lu, Huimin Cheng, Yingchuan Zhang, Shushan Wu, Rui Xie, Lin Tang, Chao Huang, Rongjie Liu, Ziyu Liu, Meizhi Yu, Yongkai Chen, Yifan Zhou, Zeliang Sun, Chang Liu, Zhen Xiang, Wei Xiao, Zixin Rao, Xinyi Liu, Yutong Hu, Mengrui Zhang, Jing Zhang, Weidi Luo, Jincheng Yu, Zhengliang Liu, Weihang You, Hanqi Jiang, Yi Pan, Junhao Chen, Xinliang Li, Tianming Liu, Wenxuan Zhong, Ping Ma

この論文をやさしく読む

ひとことで言うと

言語モデルの問題解決を、答えを長く考えることだけでなく、途中の状態を推定して検証ややり直しを選ぶ制御として整理しています。

何に役立つ?

失敗の原因に応じて、候補を増やす、検証する、巻き戻すなどの対策を選ぶための診断の枠組みになります。

この研究の面白いところ

候補を増やせば減るばらつきと、全候補に共通して残る誤りを分け、対策と原因が合っているかを問います。

どこまで分かった?

サーベイと診断的枠組みの提案です。要旨は診断仮説を新しい実験で実証した数値や、特定の制御手法の性能改善を報告していません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

大規模言語モデル(LLM)による複雑な問題解決(CPS)は、より強い推論や長い生成の問題として捉えられることが多い。しかし、初期段階の誤りの増幅、プロンプトへの脆さ、誤った判断を修正できないことは、知識や表現能力の不足だけでは説明しにくい。本サーベイは、CPSを潜在的な解の状態に関する逐次的な推定と意思決定の問題として解釈する。 制御器は、観測できない解の軌跡についての信念を保持し、雑音を含む中間的な証拠が届くたびに更新する。そして期待損失を最小化するため、判断を確定するか、検証するか、分岐するか、巻き戻すか、回答を控えるかを決める。推論は状態遷移と解釈の候補を供給し、過程の制御はそれらの提案を整え、評価し、その後の遷移と観測を調節する。 この枠組みで既存手法を、明示的な状態表現、遷移の構造化、検証と制約の実施、探索と巻き戻し、不確かさの管理という五つの要素に整理する。また、評価指標を、それが推定する統計量に応じて解釈する。さらにこの枠組みから、介入は観測された失敗に関係する誤りや不確かさの要素を狙うときに最も有効となるはずだ、という診断的な仮説が得られる。 系統的誤差、確率的誤差、不可約な誤差と、認識論的不確かさ、偶然的不確かさを区別し、この対応付けを「問題と制御の適合」、その失敗を「制御の不整合」と呼ぶ。例えば、追加のサンプリングはサンプリングのばらつきを減らせても、共通する系統誤差は変えないことがある。この視点は、現在の手法が何を推定し制御しているか、何が制御されていないか、そして信頼できる検証、対象を絞った回復、較正された不確かさ、同じ計算予算での評価が、なぜ中心的な未解決課題なのかを明確にする。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-17(UTC)
最新改訂
2026-09-17 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Complex problem solving (CPS) with large language models (LLMs) is often framed as a matter of stronger reasoning or longer generation. Yet early-step error amplification, prompt brittleness, and failures to revise incorrect commitments are difficult to explain by missing knowledge or expressive capacity alone. This survey interprets CPS as a sequential estimation-and-decision problem over a latent solution state. A controller maintains a belief about an unobserved solution trajectory, updates it as noisy intermediate evidence arrives, and decides whether to commit, verify, branch, roll back, or abstain to minimize expected loss. Reasoning supplies candidate transitions and interpretations, whereas process control shapes and evaluates those proposals and regulates subsequent transitions and observations. Within this framework, we organize existing methods around five components: explicit state representation, transition structuring, validation and constraint enforcement, search and rollback, and uncertainty management. We also interpret evaluation metrics according to the statistical quantities they estimate. The framework further yields a diagnostic hypothesis: interventions should be most effective when they target the error or uncertainty component implicated by an observed failure. We distinguish systematic, stochastic, and irreducible error together with epistemic and aleatoric uncertainty, and call this alignment problem-control fit and its failure control mismatch. For example, additional sampling may reduce sampling variability while leaving a shared systematic error unchanged. This perspective clarifies what current methods estimate and control, what remains uncontrolled, and why reliable validation, targeted recovery, calibrated uncertainty, and matched-budget evaluation are central open problems.

著者のコメント

82 pages, 7 figures. Submitted to Artificial Intelligence Review

arXiv ID: 2609.20973 / 要約の誤りについて