arXiv論文メモ
新着一覧
cs.DB · 査読状況未確認

意味に基づくデータベース検索で候補削減と高速判定を併用

Prune First, Decide Fast: Scalable Semantic Query Processing with JEVDB

Zhengle Wang and Hanxu Yan and Fuheng Zhao and Chunwei Liu

この論文をやさしく読む

ひとことで言うと

文章などの意味を使ってデータを絞り込んだり結合したりする際、すべてを大きなLLMに判断させずに高速化する仕組みです。まず候補を減らし、難しい場合だけ生成モデルへ回します。

何に役立つ?

考えられる用途は、非構造化データを含むSQL検索で遅延と推論費用を抑えることです。二つのベンチマークで速度、コスト、回答品質を評価しています。

この研究の面白いところ

通常の関係構造に対する厳密な削減と、意味的な必要条件による候補の選別を組み合わせています。候補削減と判定モデルの使い分けを別々に効かせています。

どこまで分かった?

最小遅延・最小コストという結果は評価されたクエリ群内の比較です。要旨には具体的な秒数や金額、比較システムの詳細はありません。平均F1は95.7〜97.5%であり、意味判断がすべて正しいという保証ではありません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

意味を扱うデータベースシステムは、非構造化データに対する基盤モデルの推論でSQLを拡張する。しかし現在のエンジンは、離散的な関係の判断を自己回帰型LLMに大きく依存しており、遅延と金銭的コストが高くなっている。本研究では、意味に基づくフィルター、結合、分類、順位付けに高速で型の定まった判定モデルを使い、不確かな場合だけ選択的に生成LLMへ判断を回す、拡張性の高い意味データベースシステムJEVDBを提示する。 意味結合の処理量を減らすため、JEVDBは、関係構造に対するYannakakis型の厳密な半結合による削減と、Semantic Bloom Filter(SBF)を組み合わせる。SBFは登録された必要条件を使い、潜在的な意味の辺にまたがる候補をふるいにかける。JEVDBをSemBenchと、TPC-DSから派生した意味結合ワークロードShelobで評価する。SemBenchでは、JEVDB-Flashは競争力のある回答品質を維持しながら、評価した21クエリすべてで遅延が最小となり、19クエリでコストが最小となった。 結合が54万の候補対に達するShelobでは、JEVDBは平均F1が95.7〜97.5%で、すべてのクエリを完了した。SBFによる選別は意味評価の前に候補対の87.4%を除去し、再利用可能な条件索引のスコアリングは、推論モデルへ判断を回す回数をさらに55.2%減らした。対話式クエリシミュレーター、ソースコード、ベンチマークはhttps://jevdb.orgで公開している。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-10-01(UTC)
最新改訂
2026-10-01 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Semantic database systems extend SQL with foundation-model inference over unstructured data, but current engines rely heavily on autoregressive LLMs for discrete relational decisions, creating high latency and monetary cost. We present JEVDB, a scalable semantic database system that uses fast, typed decision models for semantic filters, joins, classification, and ranking, while selectively escalating uncertain cases to generative LLMs. To reduce semantic-join work, JEVDB combines exact Yannakakis-style semijoin reduction over relational structure with Semantic Bloom Filters (SBFs), which use registered necessary conditions to screen candidates across latent semantic edges. We evaluate JEVDB on SemBench and Shelob, a TPC-DS-derived semantic-join workload. On SemBench, JEVDB-Flash achieves the lowest latency on all 21 evaluated queries and the lowest cost on 19, while maintaining competitive answer quality. On Shelob, where joins scale to 540K candidate pairs, JEVDB completes all queries with 95.7%-97.5% mean F1. SBF screening removes 87.4% of candidate pairs before semantic evaluation, and reusable condition-index scoring further reduces reasoning-model escalations by 55.2%. An interactive query simulator, source code, and benchmarks are available at https://jevdb.org.

arXiv ID: 2610.02046 / 要約の誤りについて