候補検索と順位付けを統合した大規模フィード推薦
UNIQUE: A Unified Retrieval and Ranking System for Large-Scale Feed Recommendation
この論文をやさしく読む
ひとことで言うと
推薦候補を集める処理と候補を並べる処理を一緒に学習し、情報の受け渡しの損失を減らす推薦システムです。
何に役立つ?
大規模なモバイルフィードの推薦で、検索と順位付けを効率よく連携させる用途があります。Mobile Baiduの実トラフィックで導入され、オンラインA/Bテストの改善が報告されています。
この研究の面白いところ
単層の量子化と共有表現を使い、人気の低いコンテンツも表現しやすいコード配分を狙っています。オフライン指標だけでなく視聴時間、配信量、遅延も報告しています。
どこまで分かった?
オンラインの改善値は報告されたMobile Baiduの場面での結果です。要旨にはA/Bテストの期間や標本数がなく、他サービスで同じ効果が得られるかは分かりません。P99の89 msは平均遅延ではありません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
商用のモバイルフィードシステムは、厳しい遅延制約の下で、大規模で多様かつ急速に変化するコンテンツを提供するため、候補検索と順位付けからなる処理系に依存している。しかし既存の処理系には、候補検索での階層的量子化の不安定さと、分離された検索段階・順位付け段階の間での情報損失という二つの重要な問題が残っている。これらはロングテールやコールドスタートの推薦を損ない、効率的な配信を難しくする。 本研究では、単層のフラットな量子化を用いて検索と順位付けを統合する推薦の枠組みUNIQUEを提案する。UNIQUEは、生成的なコードに基づく検索と、対象を考慮した順位付けを、一つの早期融合アーキテクチャに統合する。これにより、効率的な候補生成を維持しながら、共有表現の下でエンドツーエンドの学習を可能にする。さらに、コードブックの不均衡を緩和し、ロングテールの表現を改善するため、均衡化した量子化機構を導入する。 オフライン実験では検索と順位付けの両面からUNIQUEを評価し、コードブック分析では階層的量子化より均衡の取れた資源配分を確認した。UNIQUEをMobile Baiduのホームページフィード、発見ページ、短動画推薦の各場面に導入し、大規模な実トラフィックを処理した。オンラインA/Bテストでは、総視聴時間が0.96%、総配信量が1.08%増加し、新規利用者と非常に活動的な利用者で特に顕著な改善が得られた。配信時の測定では、P99遅延89 ms、オンライン推論のモデルFLOPs利用率(MFU)44.23%を示した。これらの結果は、UNIQUEが商用推薦における検索・順位付けの統合に対して、安定的で効率的な、実運用に対応した枠組みを提供することを示している。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-20(UTC)
- 最新改訂
- 2026-09-20 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-20 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Industrial mobile feed systems rely on a retrieval-ranking pipeline to serve large-scale, heterogeneous, and fast-changing content under strict latency constraints. However, existing pipelines still suffer from two critical issues: hierarchical quantization instability in candidate retrieval and information loss between separated retrieval and ranking stages. These issues hurt long-tail and cold-start recommendation and complicate efficient serving. To address them, we present UNIQUE, a unified retrieval and ranking recommendation framework with single-layer flat quantization. UNIQUE integrates generative code-based retrieval and target-aware ranking into one early-fusion architecture, enabling end-to-end training under a shared representation while preserving efficient candidate generation. A balanced quantization mechanism is further introduced to mitigate codebook imbalance and improve long-tail representation. Offline experiments evaluate UNIQUE from both retrieval and ranking perspectives, while codebook analysis shows more balanced resource allocation than hierarchical quantization. We deploy UNIQUE in the homepage feed, discovery-page, and short-video recommendation scenarios of Mobile Baidu, serving large-scale real-world traffic. Online A/B tests achieve a 0.96% gain in total watch duration and a 1.08% gain in total distribution volume, with notable improvements for new users and highly active users. Serving measurements show 89 ms P99 latency and 44.23% online inference MFU. These results show that UNIQUE provides a stable, efficient, and production-ready framework for unified retrieval and ranking in industrial recommendation.
arXiv ID: 2609.23718 / 要約の誤りについて