過去のニュースと予測市場で予測AIを評価する環境
Forecast-Dojo: Replayable Environments for Benchmarking and Training LLM Forecasting Agents
この論文をやさしく読む
ひとことで言うと
過去のニュースと確定済みの予測市場の問いを使い、予測AIを同じ条件で繰り返し評価・学習できる環境です。
何に役立つ?
予測エージェントの調査ツールや学習方法を比較するための課題と、結果に基づくフィードバックを提供します。
この研究の面白いところ
調査ツールは12モデルすべてのスコアを改善したものの、どのモデルも過去の市場予測には届きませんでした。
どこまで分かった?
評価は収録したPolymarketの1,568件と日付付きニュース、12モデルに基づきます。信念ノートは調査コストを下げても予測の質を一貫して上げませんでした。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
大規模言語モデルによる予測エージェントの評価と学習のため、繰り返し実行できる環境Forecast-Dojoを提案する。結果が確定済みの予測市場の問いと日付付きニュースを組み合わせ、エージェントが出来事を調査し、過去の複数の日付で予測を更新できるようにする。同じ課題とツールを使い、新たな出来事の結果が確定するのを待たずに、反復評価、学習用の対話データの収集、記録済み結果によるフィードバックを行える。Forecast-Dojoは、時系列で訓練期間と評価期間に分けたPolymarketの出来事1,568件と、日付付きニュース記事1,880万件を含む。 12モデルを評価すると、調査ツールの使用によって12モデルすべてのBrierスコアが下がった。出来事が進むにつれて予測も改善し、新しい日付の証拠がより多く記録された段階で改善幅が最大だった。それでも、全モデルがBrierスコアと正答率の両方で、過去の市場予測に及ばなかった。日付をまたいで持ち越す信念ノートは調査コストを下げたが、予測の質を一貫して改善しなかった。評価に加えて、この環境はエージェントの学習に向けた対話の軌跡と結果のフィードバックを提供し、その概念実証として教師あり微調整を示す。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-24(UTC)
- 最新改訂
- 2026-09-24 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-24 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
We introduce Forecast-Dojo, a replayable environment for benchmarking and training LLM forecasting agents. It combines resolved prediction-market questions with dated news, allowing agents to research an event and revisit their predictions at successive historical dates. The same tasks and tools support repeated evaluation, collection of training interactions, and feedback from recorded outcomes without waiting for new events to resolve. Forecast-Dojo contains 1,568 Polymarket events, split by time into training and evaluation periods, and 18.8M dated news articles. In an evaluation of 12 models, research tools lower Brier score for all 12. Forecasts also improve as events unfold, with the largest gains at steps where more newly dated evidence is recorded. Every model still trails historical market forecasts in both Brier score and accuracy. A belief notebook carried between dates lowers research cost but does not consistently improve forecast quality. Beyond evaluation, Forecast-Dojo provides interaction trajectories and outcome feedback for agent learning, with supervised fine-tuning as a proof of concept.
arXiv ID: 2609.28876 / 要約の誤りについて