arXiv論文メモ
新着一覧
cs.LG / cs.AI / stat.AP · 査読状況未確認

LLMの行動予測を補助に使い、仮説検定の標本数を減らす

Augmented Hypothesis Testing with Persona-Based LLM Simulations

Ziyad Benomar, Aymen Al Marjani, Paul Missault, Saab Mansour

この論文をやさしく読む

ひとことで言うと

AIの予測を人間の実験結果の代わりにするのではなく、必要な実験数を減らす補助情報として使う方法です。予測が外れる場合にも統計的な判断を保つ設計を目指しています。

何に役立つ?

A/Bテストで、既存の予測モデルが持つ情報を活用しながら、実験の費用や標本数を抑える用途が考えられます。ペルソナ付きLLMによる予測を使い、4データセットで評価しています。

この研究の面白いところ

効果が正か負かだけを予測する場合と、人ごとの結果を予測する場合に別の手法を用意しています。予測品質が未知でも扱えることを中心に設計しています。

どこまで分かった?

要旨には、標本数や費用の具体的削減率、定理の詳細な仮定はありません。LLMの人物シミュレーションだけで人の実験を置き換えられるという結論ではありません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

A/Bテストには、多数の標本、長い期間、大きな費用が必要である。機械学習モデルから実験結果の補助的な予測が得られる場合、その品質が不確かなため人を対象とする実験を完全には置き換えられないが、予測には有用な情報が含まれている可能性がある。本研究では、品質が未知の予測を使い、統計的妥当性を維持しながら標本数を減らす、学習支援型仮説検定の原理的な枠組みを提案する。 予測の粒度には、粗い集計レベルの情報から詳細な個人別の推定まで幅がある。本枠組みは、その両端に対応する。(1)処置効果の符号に関する二値情報しか得られない母集団レベルの方向予測に対しては、非対称な検定を用い、学習支援型アルゴリズムの枠組みで一致性と頑健性の限界を証明する。(2)個人レベルの予測に対しては、Generalized PPI++(GPPI)を導入する。これはPrediction-Powered Inferenceを拡張し、高次元変換を通して非線形な予測誤差を扱うものである。どちらの方法も、正確な予測から恩恵を受けつつ、不正確な予測や敵対的な予測に対して頑健性を保つ。 ユーザーのペルソナを与えられたAIエージェントが個人の行動を予測する、ペルソナに基づくLLMシミュレーションを使い、本枠組みを検証する。これは、両方の粒度にまたがる自然な予測源となる。4つの実世界データセットでの実験により、提案手法とペルソナに基づく予測を組み合わせると、厳密な統計的妥当性を維持しながら実験費用を大幅に削減できることを示す。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-21(UTC)
最新改訂
2026-09-21 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

A/B testing requires large sample sizes, long timelines, and significant costs. When auxiliary predictions of experimental outcomes are available from machine learning models, uncertain prediction quality precludes replacing human experiments entirely, yet these predictions may still contain useful signal. We propose a principled framework for learning-augmented hypothesis testing that leverages predictions of unknown quality to reduce sample sizes while maintaining statistical validity. Predictions naturally vary in granularity, from coarse aggregate signals to fine-grained individual-level estimates, and our framework addresses both ends of this spectrum: (1) for population-level directional predictions, where only a binary signal on the treatment effect sign is available, we use an asymmetric test and prove consistency and robustness bounds within the learning-augmented algorithms paradigm; (2) for individual-level predictions, we introduce Generalized PPI++ (GPPI), extending Prediction-Powered Inference to handle nonlinear prediction errors through higher-dimensional transformations. Both methods benefit from accurate predictions while remaining robust to inaccurate or adversarial ones. We validate our framework using persona-based LLM simulations, where AI agents equipped with user personas predict individual behavior, as a natural prediction source spanning both granularity levels. Experiments on four real-world datasets demonstrate that our methods, combined with persona-based predictions, substantially reduce experimental costs while preserving rigorous statistical validity.

著者のコメント

Work accepted at COLM Workshop on Agent Behavior

arXiv ID: 2609.24629 / 要約の誤りについて