投資ファクターを提案するAIと判定する統計手続きを分ける
Propose, Don't Judge: An Anytime-Valid Referee for LLM Agents That Mine Investment Factors
この論文をやさしく読む
ひとことで言うと
投資ファクターはAIに提案させ、採用判断は固定した統計手続きに任せる方法。
何に役立つ?
多数の候補を試す研究で、偶然よく見えたファクターの誤採用を抑える評価設計の参考になる。
この研究の面白いところ
提案者を変えると候補の産出量が変わり、判定者を変えると誤採用の数が変わると切り分けた。
どこまで分かった?
保証付きの判定では真の候補の採用も約500取引日遅れ、評価したポートフォリオのシャープレシオは審査なしより低かった。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
言語モデルを使うエージェントは、投資ファクターの提案、過去データでの検証、生き残る候補の選別、候補の廃止まで、定量的なファクター研究全体を担うようになっている。このうちどの仕事をエージェントに任せるべきかを問う。本論文の答えは、統制された自己発展である。エージェントは候補を提案できるが、判定はエージェントが変更できない固定された統計的な審査手続きが行う。審査手続きは賭けの方法を使い、候補提出後に明らかになる市場結果だけで採点するため、提案方針が何であっても、どの時点で停止しても偽発見に関する保証が成り立つ。 スクリプト、バンディット、言語モデルという三つの提案者を、この固定された審査手続きおよび意図的に情報漏れを起こす三つの審査手続きと組み合わせて比較した。評価には、正解を埋め込んだ合成環境、診断用の調査を作る環境、CSI 500を用いる10年間の逐次的な過去データ評価を用いた。誤って採用される候補の数を決めるのは審査者だった。スクリプトの提案者の下では、固定された審査手続きが基準未満のファクターを採用する数は、情報漏れのある審査手続きの5分の1から11分の1であり、どの提案者もその差を埋められなかった。一方、得られる候補の数を決めるのは提案者だった。言語モデルはスクリプトを上回り、バンディットと同程度で、バンディットにはない診断用の調査を自ら書く能力も示した。この統計的な保証には時間という代償があり、採用される真のファクターでも約500取引日待つことになり、保証付きポートフォリオのシャープレシオは審査なしの場合より低かった。判定は手続きが担い、提案と調査手段の作成はエージェントが担う。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-22(UTC)
- 最新改訂
- 2026-09-22 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-22 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Language-model agents now run the whole of quantitative factor research: they propose investment factors, backtest them, select the survivors and retire them. We ask which of those jobs an agent should keep. Our answer is governed self-evolution: the agent may propose, and a frozen statistical referee that the agent cannot touch must judge. The referee scores each candidate only on market outcomes revealed after submission, by betting, so its false-discovery guarantee holds at every stopping time for any proposal policy. We cross three proposers (a script, a bandit and a language model) with this referee and with three deliberately leaky ones, in a synthetic world with planted truth, a probe-authoring environment and a ten-year walk-forward on the CSI 500. Who judges sets the number of false admissions: the frozen referee admits 5-11 times fewer sub-threshold factors than the leaky referees under a scripted proposer, and no proposer closes that gap. Who proposes sets the yield: the language model beats the script, matches the bandit, and adds the one capability a bandit lacks, writing its own diagnostic probes. The certificate's price is time: an admitted true factor waits about 500 trading days, and the certified portfolio's Sharpe ratio therefore trails an ungated one. Judging belongs to the procedure; proposing and instrument-making belong to the agent.
著者のコメント
37 pages, 7 figures, 26 tables. JEL: G11, G14, C58
arXiv ID: 2609.27051 / 要約の誤りについて