arXiv論文メモ
新着一覧
cs.CL · 査読状況未確認

考え方を切り替えて言語モデルの推論を改善するMIRAGE

MIRAGE: Multi-Perspective Creative Language Model Reasoning with Reinforcement Learning Guidance

Arash Lagzian, Srinivas Anumasa, Dianbo Liu

この論文をやさしく読む

ひとことで言うと

問題を一つの考え方で押し通さず、代数や確率などの視点を選び、必要なら複数の視点を組み合わせる言語モデルの推論方法です。

何に役立つ?

数学・科学・論理課題で、推論時に使う解法の選択を改善する用途が考えられます。要旨では四つのベンチマークでの比較結果を報告しています。

この研究の面白いところ

視点を選ぶ役割と、選ばれた視点で解く役割を分けています。単に多くの回答を集めるだけでなく、有効な視点を優先し、必要に応じて統合する構成です。

どこまで分かった?

要旨は正解率向上と追加負荷の小ささを述べていますが、具体的な数値や実行時間は記載していません。タイトルに強化学習への言及はあるものの、その学習手順は要旨には説明されていません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

近年の大規模言語モデル(LLM)の進歩は、人工知能と、人間がAIと関わる方法を一変させた。目覚ましい進展にもかかわらず、LLMは複雑な数学的・科学的・論理的課題に苦戦している。人間の認知的柔軟性、すなわち思考の視点を動的に切り替える能力に着想を得て、本研究では推論時の創造的思考の新しい枠組みMIRAGE(Multi-perspective Inference-time Reasoning via Agent-Guided Exploration)を提案する。 MIRAGEは、有効な概念的視点、例えば代数的・確率的な視点を優先するSelectorと、確信のある解が得られるまで課題を逐次的に解き、それが得られなければ複数の視点を統合するReasonerを備える。GSM8K、MATH500、MMLU-Pro、Game-of-24のベンチマークで試験したところ、MIRAGEはChain-of-Thoughtや多様なプロンプトを用いるアンサンブルなどの方法を一貫して上回った。推論時の追加負荷を最小限に抑えつつ正解率を大きく高め、実用に向けた拡張可能な解決策を提供する。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-18(UTC)
最新改訂
2026-09-18 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Recent advances in Large Language Models (LLMs) have revolutionized artificial intelligence and how human interact with AIs. Despite impressive advancements, LLMs struggle with complex mathematical, scientific, and logical tasks. Inspired by human cognitive flexibility - our ability to dynamically switch mental perspectives - we propose MIRAGE (Multi-perspective Inference-time Reasoning via Agent-Guided Exploration), a novel inference-time creative thinking framework. MIRAGE includes a Selector that prioritizes effective conceptual perspectives (e.g., algebraic, probabilistic) and a Reasoner that sequentially solves tasks until a confident solution emerges, otherwise aggregating multiple perspectives. Tested on GSM8K, MATH500, MMLU-Pro, and Game-of-24 benchmarks, MIRAGE consistently outperforms methods like Chain-of-Thought and diverse prompting ensembles, significantly boosting accuracy with minimal inference overhead, providing a scalable solution for practical applications.

著者のコメント

18 pages, 5 figures. Accepted at the ICML 2025 Workshop on Multi-Agent Systems in the Era of Foundation Models: Opportunities, Challenges and Futures (MAS-2025)

arXiv ID: 2609.21554 / 要約の誤りについて