arXiv論文メモ
新着一覧
cs.AI / cs.IR / cs.LG · 査読状況未確認

対話AIのオフライン評価が実験結果と一致するかを検証

From Offline Proxies to Online Decisions: A Layered Engagement Evaluation Framework for Conversational AI

Xuanyi Li, Vaskar Nath, Hossein Amirkhani, Jay Li, Alex Deng

この論文をやさしく読む

ひとことで言うと

対話AIの変更を利用者に試す前に評価する指標が、実際のA/B実験の判断と合うかを調べました。

何に役立つ?

考えられる用途は、実験枠が限られるときに、試す候補を先に絞ることです。対象は著者らの配備済みアシスタントでのチェックポイント選択とプロンプト調整です。

この研究の面白いところ

固定後の8実験113組で複合指標のF1は81.1%となり、元の分類器得点の34.3%を上回りました。各予測を実験前に計算した設計です。

どこまで分かった?

主な検証は一つの配備環境における8実験113組です。他のサービスや変更種別で同じ性能になるかは要旨に記載されていません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

利用者のエンゲージメントを判断する標準的な方法はオンラインA/B実験だが、必要な利用者トラフィックと結果判定までの時間が、試せる対話AIの変更数を制限する。本研究は、変更版を利用者に見せずに計算できるオフライン指標が、実験結果と一致するかを問う。オフラインの代替指標を、行動ラベルと製品成果、学習済み分類器と候補アシスタントの振る舞い、集約したオフライン信号と実験効果という三段階の対応として捉える、再利用可能な構築・診断チェックリストを示す。併せて、点推定ではなくオフラインとオンラインの信頼区間を比較する区間を考慮した判断の一致度と、同一実験内での順位づけによって、指標全体を監査する評価手順を提案する。 評価した構成は、候補の振る舞いを採点する固定の評価セット、セッションまたはプロンプト単位のエンゲージメントを予測する分類器、サンプル単位の得点差をオンラインでのモデル単位のエンゲージメント差に写す較正層からなる。配備済みの複数ターンのアシスタントについて、モデルのチェックポイントからシステムプロンプトの調整までを含む27実験から、候補条件と対照条件の組を489件集めて監査した。主な検証には、対応づけを固定した後に実施された8実験の113組を用いた。その結果、複合指標のF1は81.1%で、元の分類器得点の34.3%を上回った。また、元の得点では方向を誤る判断が31件あったのに対し、複合指標にはなかった。過学習を防ぐため、各オフライン予測は対応する実験の開始前に計算した。これらの結果は、この配備環境で学習チェックポイントやシステムプロンプトを選ぶ際、限られた実験トラフィックを割り当てる前に候補の優先順位を決める用途を支持する。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-21(UTC)
最新改訂
2026-09-21 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Online A/B experiments are the decision standard for user engagement, but traffic and readout time limit how many conversational-AI changes can be tested. We ask whether an offline signal designed to be computable without treatment-arm user exposure agrees with the outcomes of those experiments. We contribute a reusable construction and diagnosis checklist that treats an offline proxy as a chain of three alignments: behavioral label to product outcome, learned classifier to candidate-assistant behavior, and aggregated offline signal to experiment effect. A companion evaluation protocol audits the whole composite by interval-aware decision agreement, which compares offline and online confidence intervals instead of point estimates, and by within-experiment ranking. The instantiation we evaluate comprises a fixed evaluation suite on which candidate behavior is scored, an engagement classifier trained to predict session/prompt level engagements, and a calibration layer mapping sample-level score differences to online model-level engagement deltas. We then report the audit: 489 paired offline-online contrasts (one candidate arm against its control) from 27 experiments on a deployed multi-turn assistant, spanning model checkpoints to system-prompt tuning. Our primary test uses the 113 contrasts from eight experiments that ran after the map was frozen: on these the composite reaches 81.1% F1, against 34.3% for the raw classifier score it is built on, and makes no wrong-direction calls where that raw score makes 31. Every offline prediction was computed before its experiment ran to prevent overfitting. The evidence supports using the composite to prioritize candidates before scarce experiment traffic is allocated---in our deployment of the experiment, selecting among training checkpoints and tuning system prompts.

著者のコメント

11 pages main text, 9 pages supplementary material; 2 figures, 25 tables

arXiv ID: 2609.25408 / 要約の誤りについて