対話型AIはいつ検索し、何を引用して答えるのか
Characterizing Web Search by Conversational LLM Agents: From Search Decisions and Strategies to Results and Responses
この論文をやさしく読む
ひとことで言うと
対話AIが検索するかどうかを決める段階から、検索語を作り、結果を選び、引用して答えるまでを比較しています。
何に役立つ?
検索回数だけでAIの回答品質を測らず、検索判断、情報源、引用の対応を含めて評価・設計するための材料になります。
この研究の面白いところ
実際のユーザー利用と統制実験を組み合わせています。検索結果を使っていても、その結果が引用されず、根拠の追跡が難しくなる場合を指摘しています。
どこまで分かった?
4プラットフォームの比較であり、要旨には対象モデルの詳細や観測期間、定量値は記載されていません。検索頻度と品質の関係も、頻度を増やせば必ず悪くなるという意味ではありません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
対話型LLMエージェントはウェブ検索への依存を強めているが、エージェントによる検索の最初から最後までの流れは、まだよく理解されていない。本研究では、4つの主要な対話プラットフォーム、ChatGPT、Claude、Grok、DeepSeekを横断したウェブ検索の初の研究を提示する。実際のユーザーとのやり取り(in vivo)と、同じプラットフォームのモデルをAPI経由で使う統制実験(in vitro)を組み合わせる。 ウェブ検索を実行する判断の質、クエリを組み立てる戦略、受け取る検索結果におけるドメイン選好の可能性、そして検索結果を根拠のある回答へ変換する際の選択を調べる。ウェブ検索の判断はプラットフォームやモデルによって大きく異なり、検索を頻繁に行うことが必ずしも回答の質の向上につながらないことが分かった。 さらに、対話エージェントはそれぞれ異なる複雑なクエリ戦略を使い、プラットフォーム固有の検索エンジンは、それぞれが優先するドメインの検索結果を返すことを示す。最後に、回答は大部分が検索結果に根拠を置いているものの、一部の主張は引用されていない検索結果に依拠しており、出典の明示と信頼性に懸念が生じる。本研究の知見は、対話型の情報取得に最適化した将来のAIエージェントとウェブ検索ツールの設計に重要な意味を持つ。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-16(UTC)
- 最新改訂
- 2026-09-16 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-16 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Conversational LLM agents increasingly rely on Web search, yet the end-to-end lifecycle of agentic search remains poorly understood. We present the first study of Web search across four major conversational platforms (ChatGPT, Claude, Grok, and DeepSeek), combining real-world user interactions (invivo) with controlled experiments using the same platform's models by their APIs (invitro). We investigate the quality of agentic decisions to invoke Web search, their strategies to formulate queries, the potential domain preferences in the search results they receive, and the choices they make when transforming search results into grounded responses. We find that Web-search decisions vary substantially across platforms and models, while more frequent Web-search invocation does not necessarily yield better response quality. We further show that conversational agents employ different complex querying strategies and that platform specific search engines return search results from their preferred domains. Finally, although responses are largely grounded in search results, some claims rely on uncited search results, raising concerns about attribution and reliability. Our findings have important implications for the design of future AI agents and Web search tools optimized for conversational retrieval.
arXiv ID: 2609.19244 / 要約の誤りについて