33種類の文書検索モデルを複数分野で比較
A Systematic Multi-Domain Evaluation of Document Retrievers
この論文をやさしく読む
ひとことで言うと
33種類の文書検索モデルを、7つのデータセットで品質と速さの両面から比べた。
何に役立つ?
文書検索器を選ぶとき、精度と問い合わせ遅延の兼ね合いを検討する材料になる。
この研究の面白いところ
同じ計算予算と公開設定で比較し、NV-Embed-v2の強い精度と大きな遅延、SPLADE-v3の速さと精度を示した。
どこまで分かった?
モデルごとの個別調整はしていない。結論は対象の7データセットと再現した公開設定に基づく。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
文書検索は現代のAIシステムの多くで重要であり、後段の課題の有効性、頑健性、公平性に直接影響する。近年、検索モデルは増えたが、文献の比較研究は対象が狭かったり、単一のベンチマーク、分野、モデル群に集中したりすることが多い。そのため、検索モデルどうしの強み、弱み、トレードオフについて信頼できる結論を出しにくい。著者らはこの不足に対し、疎な手法、密な手法、拡張に基づく手法という3系統の33モデルを、7つの情報検索データセットで大規模に実証評価し、検索品質、実行時間、失敗しやすい箇所を分析する。モデルごとに個別の調整はせず、公開文書から再現できる設定と同一の計算予算で、そのまま評価する。結果ではNV-Embed-v2が7データセット中4つで最高の性能を示す一方、問い合わせ時の遅延が大きい。疎な検索モデルではSPLADE-v3が遅延を大きく抑えながら最高水準の手法に匹敵し、MS MARCOでは最高得点を得た。指示追従型のデータセットではGritLMが最良だった。失敗箇所の分析からはモデルや系統間の違いが見え、検索性能にはまだ活用されていない改善の余地があると分かった。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-24(UTC)
- 最新改訂
- 2026-09-24 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-24 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Document retrieval is a crucial component of many modern AI systems, directly influencing their effectiveness, robustness, and fairness in downstream tasks. While recent years have seen a growing number of retrievers, comparative studies in the literature are typically limited in scope or focused on singular benchmarks, domains, or model families. This fragmentation makes it difficult to draw reliable conclusions about the relative strengths, weaknesses, and trade-offs of document retrievers. To address this gap, we conduct a large-scale empirical evaluation of document retrievers, covering three families (sparse, dense, and expansion-based) and evaluating 33 retrievers across seven IR datasets, analyzing retrieval quality, runtime, and failure points. Rather than tuning each model individually, we evaluate every retriever off the shelf, under the configuration reconstructable from its public documentation and a uniform compute budget. Our results show that NV-Embed-v2 achieves the strongest performance on four of the seven datasets, albeit at the cost of substantial query latencies. Among sparse retrievers, we find that SPLADE-v3 rivals the top-performing approach despite much lower latency, and even achieves top scores on MS MARCO. On instruction-following datasets, GritLM delivers the best performance. Finally, an analysis of the retrievers' failure points reveals contrasts between models and families that indicate potential for unrealized gains in retrieval performance.
arXiv ID: 2609.29455 / 要約の誤りについて