arXiv論文メモ
新着一覧
cs.CL / cs.AI · 査読状況未確認

営業担当者の楽観的な主張がAIの商談判定を誤らせる

Persuaded, Not Informed: Incentive-Misaligned Witnesses Defeat In-Context Grounding

Rahul Balakavi

この論文をやさしく読む

ひとことで言うと

営業記録の中にある利害関係者の楽観的な発言を、AIが会社の正式な条件より強い証拠として扱う失敗を調べています。

何に役立つ?

考えられる用途は、CRM記録を読むAIの評価方法や、価格表・設置方針に照らす判定手順の設計です。要旨が示すのは評価課題での誤判定傾向です。

この研究の面白いところ

情報不足と説得の影響を分け、同じ情報を与える対照実験や予算・納期の計算担当だけを変える実験で失敗要因を調べています。

どこまで分かった?

最も強いモデルでは比較した条件の差が信頼区間内であり、性能の下限は証明していません。事前に定めた一般化テストも否定的で、現象の前提は方針が入力に正確に示されていることです。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

言語モデルを使うエージェントは、営業見込み客を適格とみなすかなど、顧客関係管理(CRM)の記録に基づく質問に答える機会が増えている。本研究は、より強いモデルを使うだけでは解消しない失敗の形を特定する。記録内に楽観的に述べる動機のある当事者、ここではCRMに発言が記録された営業担当者の主張があると、モデルはそれを証拠として扱い、会社自身の記録では受け入れられない商談を承認してしまう。 CRMArena-Proの見込み客判定100件では、営業担当者はすべての通話で許容される納期を、76件で許容される予算を主張した。こうした主張が価格表と設置方針に反する31件について、会話記録だけを読んだモデルは31件中29件で商談を承認した。4事業者の7モデルでも同じ傾向が見られ、87~97%で誤誘導された。モデルの規模や明示的な推論による耐性は見られなかった。実際の失敗35件のうち、主張がない例は3件だけだった。著者らは、問題の中心を情報不足ではなく説得の影響とみなす。 提案するのは新しいモデル構成ではなく診断法である。第一に、説得の影響と情報の欠落を分ける分類分析を行う。第二に、同じ情報を与える対照実験では、モデルに記録を渡すと厳密な正解率が41から18へ下がり、再現率は上がる一方で適合率が大きく低下した。第三に、情報抽出を固定して、予算と納期を誰が計算するかだけを変える対照実験を行った。差は低価格帯のモデルでは42ポイント、すでに正しく計算するモデルでは2~5ポイントだった。最も強いモデルでは条件間の結果が信頼区間内に収まるため、この傾向は一貫した方向と健全性に関する性質を示すが、性能の下限を証明したものではない。事前に定めた一般化テストは否定的な結果となった。この問題が生じる前提として入力内に方針が正確に明記されていることを整理し、評価資料をすべて公開している。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-23(UTC)
最新改訂
2026-09-23 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Language-model agents increasingly answer questions over customer-relationship management (CRM) records, such as whether to qualify a sales lead. We identify a failure mode not addressed by a stronger model: when the context contains an assertion by a party with an incentive toward optimism - here the sales representative, a witness recorded in the CRM - the model treats the assertion as evidence and clears deals the company's own records deem unacceptable. Across 100 lead-qualification tasks from CRMArena-Pro, the representative asserts an acceptable timeline in every call and an acceptable budget in 76; on the 31 tasks where such an assertion contradicts the price list and installation policy, a model reading only the transcript clears the deal in 29 of 31 cases. The signature is consistent across seven models from four providers (misled on 87-97%); scale and explicit reasoning confer no resistance. Only 3 of 35 genuine failures involve no assertion: the failure is persuasion, not missing information. We contribute a diagnostic method rather than an architecture: (i) a bucket analysis that separates persuasion from information gaps, (ii) a same-information control showing that supplying the records to the model lowers strict accuracy from 41 to 18 while raising recall - precision collapses - and (iii) a compute-step control that holds extraction fixed and varies only who computes Budget and Timeline. The margin ranges from 42 points on an inexpensive model to 2-5 points on models that already compute correctly; on the strongest models the arms are within confidence intervals, so the pattern is a consistent direction and a soundness property, not a proved performance floor. We pre-specify a generalization test that returns a negative result, characterize the precondition (a policy exactly specified in the inputs), and release all evaluation artifacts.

著者のコメント

9 pages, 4 figures, IEEE conference format. Ancillary files contain the evaluation harness, pre-specifications, and per-run result files

arXiv ID: 2609.28854 / 要約の誤りについて