予測ツールの費用を含めて言語モデルの意思決定を評価
Forecast Workflow Bench: Evaluating Language-Model Decisions with Budgeted Forecast Tools
この論文をやさしく読む
ひとことで言うと
言語モデルが予測ツールにいくら使い、得た予測を電力やレンタサイクルの供給量決定に生かせるかを測るベンチマークです。
何に役立つ?
予測の正確さだけでなく、予測を買う費用とその後の意思決定の損失を合わせて、エージェントを評価できる。
この研究の面白いところ
GPT-6 Astraは予算の2.5%だけで短期予測を選択的に使い、三通りの採点条件で固定方針を上回った。
どこまで分かった?
結果は1251件の電力・レンタサイクル事例と模擬的な契約での評価であり、実際の契約運用での成果は要旨に示されていない。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
時系列の基盤モデルは業務上の意思決定に使う予測を提供するが、予測精度だけでは価値を決められない。これらのモデルを使うエージェントの評価には、意思決定の質と予測の費用を測る必要がある。FWBenchは、固定された予測ツールと模擬的な供給能力の契約を使い、電力とレンタサイクルの1251事例でこの能力を評価する。エージェントはモデル、過去データの範囲、予測期間を選んだ後、明示された損失と費用の目的関数を最小化するよう供給能力を提出する。二つのホスト型と八つのローカル環境設定を評価し、小型の言語モデルも含めた。ローカルモデルでは、時系列基盤モデルを使う場合と使わない場合の両方を試した。GPT-6 Astraは安価な短期予測を選択的に購入し、予算の2.5%だけを使った。保存された意思決定を、損失と費用に対する三通りの重み付けで採点すると、固定方針を上回った。FWBenchは、言語モデルが費用の制約の下で時系列予測をどう選び、意思決定に使うかを再現可能な形で評価できる。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-23(UTC)
- 最新改訂
- 2026-09-23 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-23 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Time-series foundation models (TSFMs) provide forecasts for operational decisions, but accuracy alone does not determine their value. Evaluating agents that use these models requires measuring decision quality and forecast cost. FWBench evaluates this capability on 1,251 electricity and cycle-hire cases using fixed forecast tools and simulated capacity contracts. Agents select models, histories and horizons, then submit capacities to minimize a stated loss-cost objective. We evaluated two hosted and eight local configurations, including small language models, and tested local models with and without TSFMs. GPT-6 Astra bought inexpensive short-horizon forecasts selectively, using 2.5% of the budget, and outperformed fixed policies when the saved decisions were scored with three loss-cost weightings. FWBench enables reproducible evaluation of how language models select and use time-series forecasts to make decisions under cost constraints.
arXiv ID: 2609.27385 / 要約の誤りについて