LLMの品質を満たす最安費用を競争入札で学ぶ
Learning the Cost of Reliable Inference
この論文をやさしく読む
ひとことで言うと
LLMの問い合わせごとに、必要な品質を満たす提供者を入札で選び、費用を抑える仕組みを提案した。
何に役立つ?
複数のモデル提供者から品質条件を守りながら選ぶ、料金設計と経路選択の検討に役立つ。
この研究の面白いところ
提供者の品質を実際に振り分けながら学び、正直な費用見積もりを促す第二価格オークションを組み合わせた。
どこまで分かった?
LlamaとQwenの数学・質問応答ベンチマークによる実験であり、節約は競争的な市場条件がある場合の可能性として述べられている。
v2のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
ベンチマークと経路選択のプラットフォームは、大規模言語モデルの提供者と利用者をつなぐ仲介役になりつつある。しかし、提供者は通常トークンごとの固定価格を使うため、利用者は自分の課題に対して最も競争力のある価格を得にくい。本研究は、課題ごとのトークン価格を提供者間の競争で決め、保証された品質水準を満たす中で競争的な価格を得られる調達プラットフォームを設計する。問い合わせを逆向きの第二価格オークションで順次振り分け、提供者がその問い合わせを処理する平均費用の最善の見積もりを正直に入札するよう促す。振り分けながら各提供者の品質を学び、希望する品質しきい値を満たす提供者の中で、費用面で最も競争力のあるところへ徐々に割り当てる。 設計を検証するため、LlamaとQwen系列の複数のモデルを使い、一般的な数学的推論と質問応答のベンチマークで実験した。その結果、プラットフォーム上で最も費用競争力のある提供者の価格マージンは、課題と品質しきい値によって10%から71%まで大きく変わった。これは現在の固定価格市場に相当の非効率があることを示唆し、競争的な市場条件が整う場合、プラットフォームによって利用者が最大限の節約を得られる可能性を示す。
v2の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-23(UTC)
- 最新改訂
- 2026-09-24 · v2
- 査読・掲載
- 査読状況未確認
更新履歴
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Benchmarking and routing platforms increasingly act as intermediaries connecting large language model providers with end-users. However, providers on these platforms typically use a fixed price per token, preventing users from achieving the most competitive price for their tasks. In this work, we design a procurement platform where token prices for each task are driven by provider competition, enabling users to secure competitive pricing for guaranteed quality levels. To this end, the platform sequentially routes queries via a reverse second-price auction that incentivizes model providers to truthfully bid their best estimate of the average cost to serve a user's query. As it routes queries, the platform learns the quality offered by each provider and progressively routes queries to the most cost-competitive provider among those meeting a desired quality threshold. To validate our design, we conduct experiments with multiple LLMs from the Llama and Qwen families on popular mathematical reasoning and question-answering benchmarks. The results show that the pricing margin of the most cost-competitive provider on our platform varies significantly---from $10\%$ to $71\%$---depending on the task and quality threshold. This suggests a substantial inefficiency in the current fixed-price market, and it demonstrates that our platform may enable users to capture maximum savings whenever competitive market conditions permit.
arXiv ID: 2609.28322 / 要約の誤りについて