arXiv論文メモ
新着一覧
cs.AI / cs.CL · 査読状況未確認

顧客対応AIを本番投入前に模擬対話で評価

Screen Before You Serve: Simulation for Production Customer Experience AI Agents at 140M Scale

Edesio Alcoba, Kevin Rossell, Aman Gupta, Shao Tang, Jiwoo Hong, Pabel Carrillo-Mendoza, Wanderson Conceição Ferreira, Alvaro Tedeschi, Zayd Simjee, Shreya Rajpal, Bruno Finardi Hime, Christian Sousa, Luis Moneda, Herbert Fei, Daniel Silva, Rohan Ramanath

この論文をやさしく読む

ひとことで言うと

顧客対応AIを公開する前に、模擬顧客と模擬ツールで試して候補を選ぶ方法です。

何に役立つ?

顧客へ影響を与えずにモデルや設定を比較し、本番導入前のエージェント評価を広げる際に役立ちます。

この研究の面白いところ

シミュレーションの評価が4つの配備版の本番評価と高く相関し、本番A/BテストでもtNPSやセルフサービス率の改善が示されました。

どこまで分かった?

本番での改善はNubankの対象エージェントでの結果です。後半のA/Bテストではセルフサービス率は上がりましたが、tNPSの有意な変化はありませんでした。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

顧客対応(CX)エージェントは、ツールと大規模言語モデルを使って顧客の依頼に応え、組織の製品について対話を進める。特に規制業種では、意図の検出、複雑な運用方針の順守、ツールの確実な利用が必要で、改善が難しい。手動の一連のテストは網羅性に限界があり、本番の実験では失敗を顧客に経験させて信頼を損なう可能性がある。本研究は、候補となるCXエージェントを配備前にふるいにかける、仮説に基づいたシミュレーションの作業手順を提示する。合成された顧客はエージェントの応答に反応し、模擬ツール出力を用いることで本番の基幹システムを呼ばずに複数段階の処理を試せる。Snowglobeシミュレーターを、Nubankのカード配送エージェントと、その後継でブラジルにおける同社最大件数のチャットサポートであるカード管理エージェントに適用した。配備された4バージョンについて、模擬環境と本番環境でのバージョン単位の二値評価スコアには高い相関があった。シミュレーションを使った反復改善の結果、本番のA/Bテストでは取引関連のネットプロモータースコア(tNPS)が36.69ポイント上がった。また、16,000件超の模擬対話で公開重みモデルの設定を選別した。その後の本番A/Bテストでは、選んだモデルによってセルフサービス率(SSR)が8.82ポイント上がり、Nubankで観測された最高水準に達した一方、tNPSに統計的に有意な変化はなかった。シミュレーションにより、顧客を実験にさらさずモデル、推論設定、プロンプトを広く試せるようになり、本番実験だけでは実行が難しかった改善を可能にした。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-24(UTC)
最新改訂
2026-09-24 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Customer experience (CX) agents use tools and large language models to address customer requests and guide conversational interactions with an organization's products. Improving these agents, especially in regulated industries, is difficult: they must detect intent, follow complex operational policies and use tools reliably. Manual end-to-end testing offers limited coverage, while live experiments expose customers to failures that can erode trust. We present a hypothesis-driven simulation workflow for screening candidate CX agents before deployment. Synthetic customers react to agent responses and simulated tool outputs enable multi-step agentic workflows without invoking production backends. We use the Snowglobe simulator on Nubank's Card Delivery agent and its expanded successor, Card Management - Nubank's highest-volume chat-support agent in Brazil. Across 4 deployed versions, simulated and production version-level binary evaluator scores show high correlation. Simulation-guided iteration increased transactional net promoter score (tNPS) by 36.69 points in a live A/B test. We also screened open-weight configurations in over 16,000 simulated conversations. In a subsequent live A/B test, the selected model increased self-service rate (SSR) by 8.82 percentage points to the highest level observed at Nubank, with no statistically significant change in tNPS. Simulation made broad exploration of models, reasoning settings, and prompts feasible without customer exposure, enabling production improvements that would have been impractical to pursue through live experimentation alone.

著者のコメント

17 pages, 11 figures

arXiv ID: 2609.30137 / 要約の誤りについて