arXiv論文メモ
新着一覧
econ.GN / q-fin.EC · 査読状況未確認

将来データを使った学習は金融予測を本当に良くするか

Does Training on Future Data Pay? Look-Ahead Bias in Forecasting with Pretrained Models

Haiqiang Chen, Li Chen, Yunlong Chen, Difang Huang, Bo Zhang

この論文をやさしく読む

ひとことで言うと

予測時点より未来の情報で事前学習した金融モデルが、評価上必ず有利になるのかを比較します。

何に役立つ?

時系列モデルの評価で、情報漏洩の有無と、精度や投資価値への実際の影響を分けて考えるための研究です。14市場と四つの予測期間を扱います。

この研究の面白いところ

同じ数値履歴・推論手順で学習年次だけを変えます。米国学習環境では20組中18組で平均二乗誤差が増え、将来情報を含む学習がむしろ精度を下げる例を示します。

どこまで分かった?

全世界や因子を加えた学習では結果が混在します。将来情報の利用が許されるという結論ではなく、情報集合の違反だけでは性能の過大評価を断定できないという区別です。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

予測起点より後の情報を学習に使うことで、金融予測の測定上の精度と経済価値が過大評価されるかを調べる。金融時系列の基盤モデル5セットを、14の株式市場と4つの予測期間にわたって評価する。各セットは、米国、世界、ファクターを追加した学習環境のもとで、それぞれ独立に学習した年次版からなる。ローリング比較では、予測対象を固定してモデルの年次版を変える。固定版比較では、年次版を固定し、その学習期限をまたぐように予測対象期間を動かす。それぞれの代替予測を、同じ数値履歴と推論手順を用いた、予測起点に整合する時点情報(PIT)ベンチマークと対にして比較する。 米国で学習した基準環境では、起点後の情報を含む版は、予測力のあるPIT予測を大きく修正するものの、両方の比較設計で概して精度を低下させる。ローリング比較をまとめると、米国におけるモデルセットと予測期間の20の組み合わせのうち18で、平均二乗予測誤差が増える。予測起点をまたぐ更新も、同じ長さの起点前の更新に比べ、平均的に成績が悪い。1か月先予測に共通の制約付き資産配分ルールを適用した場合、年率の確実性等価収益率について、起点後情報を含む予測からPIT予測を引いた差の中央値は、米国で−1.77パーセントポイント、国際市場で−2.14ポイントとなる。世界を対象にした学習やファクターを追加した学習では、予測への影響はよりまちまちである。 二乗誤差の厳密な分解から、予測修正が精度を改善するのは、誤差を訂正する便益が修正量の平均二乗を上回る場合であることが分かる。米国での学習では、修正とPIT誤差の整合性は概してこの条件を満たさない。したがって、将来時点の情報への接触は利用可能情報集合への違反を示すが、予測精度や投資家にとっての価値が過大評価されたことの十分な証拠にはならない。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-17(UTC)
最新改訂
2026-09-17 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

We examine whether post-origin training information inflates the measured accuracy and economic value of financial forecasts. We evaluate five sets of financial time-series foundation models, each comprising independently trained annual vintages under U.S., global, and factor-augmented training environments, across 14 equity markets and four forecast horizons. Rolling comparisons vary the annual vintage for a fixed forecast; fixed-vintage comparisons hold the vintage fixed as target windows move across its training cutoff. Each alternative forecast is paired with an origin-aligned point-in-time (PIT) benchmark using identical numerical histories and inference protocols. In the U.S.-trained reference environment, post-origin vintages materially revise informative PIT forecasts but generally reduce accuracy in both designs. Pooled rolling comparisons yield higher mean squared forecast errors in 18 of 20 U.S. model-set-horizon combinations. The origin-crossing update also performs worse on average than an equally long pre-origin update. Under a common constrained allocation rule using one-month forecasts, median exposed-minus-PIT differences in annualized certainty-equivalent returns are -1.77 percentage points in the United States and -2.14 points internationally. Global and factor-augmented training produce more mixed predictive effects. An exact squared-error decomposition shows that revisions improve accuracy when their error-correcting benefit exceeds their mean squared magnitude; under U.S. training, alignment with PIT errors generally falls short of this requirement. Temporal exposure therefore establishes an information-set violation, not sufficient evidence of inflated predictive accuracy or investor value.

arXiv ID: 2609.20554 / 要約の誤りについて