情報を自分で検索する予測エージェントを結果報酬で学習
Do Your Own Research: Learning to Forecast by Learning to Search
この論文をやさしく読む
ひとことで言うと
出来事を予測するAIに、答えだけでなく、どの情報を検索して読むかも結果の良し悪しから学ばせる研究です。
何に役立つ?
時点を区切った情報だけを使う予測エージェントの学習と評価に役立ちます。公開される環境と記録は、検索行動と予測性能を合わせて比較するための基盤になります。
この研究の面白いところ
学習後は検索回数が減る一方で確率予測の較正が改善したと報告しています。用意された資料から予測するだけでなく、資料を獲得する過程まで報酬の対象にしています。
どこまで分かった?
評価は過去に結果が確定した質問を使い、締切前の情報へ制限したものです。将来の現実の予測で同じ優位性を保証する結果ではありません。示されたsoft-Brier差は0.002で、要旨にはその統計的不確かさはありません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
結果に基づく強化学習は、言語モデルに現実の出来事を予測する能力を学習させられる。しかし従来の予測研究では、調査用の文脈を学習前に固定するか、エージェントによる調査をテスト時にだけ使っており、証拠を集める技能が報酬によって形成されることはない。 本研究では、結果が確定したPolymarketの質問2100件以上から構築した、エージェント型予測の環境、データセット、実行・評価基盤を導入する。エージェントはロールアウト時に、ウェブ検索、ページ閲覧、金融時系列を通じて自ら文脈を獲得する。これらはすべて、重層的な漏洩フィルタによって、各質問の締切前に公開された情報に制限される。この環境上で、Qwen3.5-35B-A3B(有効パラメータ30億)を、ブライアスコアの報酬と1エポックのGRPOで学習する。 学習により、エージェントの情報との関わり方が変わる。確率予測の較正は30〜40%改善し、証拠の扱いを学ぶにつれ、ロールアウトごとの検索試行回数は3.8回から2.25回へ減る。同一の実行・評価基盤で4つの最先端モデルと比較すると、学習済み方策は、証拠に基づく予測で評価したすべての最先端モデルを上回る。その中にはClaude Opus 4.5も含まれ、soft-Brierは0.254対0.256、n=265である。推論コストは約5%であり、差が最も大きいのは、群衆自身も判断が定まっていなかった最も難しい質問である。時間に関する予測エージェント用の再利用可能な実行・評価基盤として、環境、データセット、ロールアウトごとの記録を公開する。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-10-01(UTC)
- 最新改訂
- 2026-10-01 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-10-01 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Outcome-based reinforcement learning can train language models to forecast real-world events, but prior forecasting work either freezes research context before training or deploys agentic research only at test time, so the skill of gathering evidence is never shaped by the reward. We introduce an agentic forecasting environment, dataset, and harness built from 2,100+ resolved Polymarket questions; the agent acquires its own context at rollout time (web search, page reading, and financial time series, all restricted by layered leak filtering to information published before each question's cutoff), and we train Qwen3.5-35B-A3B (3B active parameters) on it with single-epoch GRPO under a Brier-score reward. Training changes how the agent interacts with information: calibration improves 30-40%, and search attempts fall from 3.8 to 2.25 per rollout as evidence discipline is learned. Evaluated in an identical harness against four frontier models, the trained policy also finishes ahead of every frontier model tested at evidence-based forecasting, including Claude Opus 4.5 (soft-Brier 0.254 vs. 0.256, n=265), at about 5% of the inference cost, and its margin is widest on the hardest questions, the ones the crowd itself had not decided. We release the environment, dataset, and per-rollout records as a reusable harness for temporal forecasting agents.
著者のコメント
Accepted at the NeurIPS 2026 Workshop on Foundation Models for Temporal Systems (FMTS). 9 pages, 4 figures. Code and data: https://github.com/afifi-yusuf/prime-forecast
arXiv ID: 2610.01955 / 要約の誤りについて