arXiv論文メモ
新着一覧
cs.IR / cs.AI · 査読状況未確認

国をまたぐ商品推薦で購買履歴を条件付きで混合

Cross-Country Code-Mixing for Generative Recommendation

Yuan Gao, Hao Deng, Haibo Xing, Yi Xu, Lingyu Mu, Jinxin Hu, Yu Zhang, and Xiaoyi Zeng

この論文をやさしく読む

ひとことで言うと

国ごとに商品と利用者のIDが違う推薦システムで、条件の合う商品を履歴の中で置き換えて学習します。

何に役立つ?

購買データの少ない市場の推薦を改善する用途が考えられます。要旨は実データとオンラインA/Bテストでの改善を報告しています。

この研究の面白いところ

商品内容だけでなく価格、利用者層、人気も見て国をまたぐ置換を行い、混合した履歴の自然さに応じて学習時の重みを調整します。

どこまで分かった?

報告された収益と注文数の増加は評価した大規模プラットフォームでの結果です。ほかの市場でも同じ増加率になるとは示されていません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

現代の電子商取引では、国をまたぐ推薦システムでも市場ごとに利用者IDと商品IDの空間が分かれていることが多く、従来の領域横断手法が頼る共通の対応点が失われる。生成的推薦は商品を共通のトークン空間へ写し、一つのモデルを学習することでこの問題を和らげる。しかし、既存の方法は行動履歴を国ごとに厳密に分けるため、知識の移転はモデルのパラメータだけで起こり、データの段階では起こらない。多言語の自然言語処理におけるコードスイッチングのデータに着想を得て、本研究は、二種類の制約と文脈を考慮したコード混合によって、データ段階で国をまたぐ教師信号を加える生成的推薦の枠組みCMRecを提案する。 まず、複数国の商品について、複数の種類のコンテンツと行動の共起から共通の意味コードブックを学ぶ。次に、コンテンツに関する静的制約と、価格、利用者層、人気などの動的制約の両方を満たすトークン単位の置換により、複数国の商品が混じる系列を合成する。最後に、現在の系列における自然さに応じて混合標本の重みを変える、文脈を考慮した損失を導入する。二つの実際の複数国データセットとオンラインA/Bテストでは、データの少ない国の推薦品質を大きく改善しつつ、データの多い国の性能を維持した。大規模な電子商取引プラットフォームでは、広告収益が1.77%、注文数が2.64%増えた。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-24(UTC)
最新改訂
2026-09-24 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Cross-country recommendation on modern e-commerce platforms is typically deployed with disjoint user and item ID spaces across markets, removing the shared anchors that conventional cross-domain methods rely on. Generative recommendation (GR) mitigates this by mapping items into a shared token space and training a unified model, but existing approaches keep behavior sequences strictly country-specific, so knowledge transfer occurs only at the parameter level and remains absent at the data level. Inspired by code-switching corpora in multilingual natural language processing, we propose CMRec, a cross-country GR framework that injects cross-country supervision at the data level via dual-constrained, context-aware code-mixing. CMRec first learns a shared semantic codebook from multi-modal content and behavioral co-occurrence across countries. It then uses this codebook to synthesize mixed-country sequences via token-level substitutions that satisfy both static (content) and dynamic (e.g., price, audience, popularity) constraints. Finally, it introduces a context-aware loss that reweights mixed samples according to their plausibility in the current sequence. Experiments on two real-world multi-country datasets and an online A/B test show that CMRec substantially improves recommendation quality in data-sparse countries while preserving performance in data-rich countries, achieving +1.77% advertising revenue and +2.64% orders on a large-scale e-commerce platform.

著者のコメント

CIKM 2026 Short

arXiv ID: 2609.28972 / 要約の誤りについて