arXiv論文メモ
新着一覧
cs.AI / cs.IR · 査読状況未確認

推薦モデルの研究を実験結果から次の課題へつなぐAgentX-Model

Advancing Model Research in AgentX: Long-Horizon Autonomy for Industrial Recommender Systems

Shuang Yang, Zijie Zhuang, Changxin Lao, Pengbo Xu, Hanwen Xu, Yusheng Huang, Han Gao, Guanchen Wang, Tianbao Ma, Linxun Chen, Peilin Song, Xuming Wang, Chen Li, Fan Wu, Tao Wang, Zibo Zhao, Xiangyu Wu, An Liu, Fei Pan, Peng Jiang, Chen Yang, Zhaojie Liu, Wenwu Ou

この論文をやさしく読む

ひとことで言うと

推薦モデルの実験結果を次の研究課題につなぐ二つのエージェントの仕組みを示し、本番環境で評価した研究。

何に役立つ?

推薦システムの継続的なモデル改善において、提案、実験、診断、次の実装選択をどうつなぐかの参考になる。報告された事業指標は評価した環境での結果である。

この研究の面白いところ

モデル変更実験636件のうち560件でAUCが事業上の基準を超えた一方、候補を具体的に分析・選択しているなら、複雑な実験スケジューリングは一貫した効率向上を示さなかった。

どこまで分かった?

オンラインA/B評価は異なる事業環境での最新五件についての報告であり、示した改善幅が他の推薦システムにも当てはまるとは要旨に記載されていない。スケジューリングに関する結果は初期結果とされる。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

産業用の推薦システム研究を継続するには、一つの実験結果を使って次に調べることを決める必要がある。本研究は、AgentXの次世代モデル研究枠組みであるAgentX-Modelを提示する。事業側の入力と予測課題で定めたサンドボックス内で、提案の作成とモデル実験を結び付ける。AgentX-Modelは研究エージェントとモデルエージェントからなる二つのエージェント構成を採用する。研究エージェントは論文や実験結果から独立したレビューを受ける提案を作り、モデルエージェントは複数回にわたる調査を実施して、コード、測定値、未解決の問いを返す。研究エージェントは返された結果を使って出発点となる実装を選び、次の研究課題を定める。これにより、後続の実験を先行する知見の上に積み重ねられる。 この継続的な研究を、再現、追試、組み合わせ、診断という四つの行動に整理する。最初の三つは通常の研究を進め、診断は修正を選ぶために必要な証拠を集める。後者には、PCOCで測定する予測バイアスなど、事業側のフィードバックやオンライン評価から指摘された問題も含まれる。本番環境での評価では、完了したモデル変更実験636件のうち560件で、AUCが事業側の基準値を上回った。研究の継続に伴い、一部の実験では同じ系譜内の比較可能なすべての先行実験を上回るAUCを記録した。異なる事業環境での最新五件のオンラインA/B評価では、顧客獲得効率の10~15%、対象層向け広告支出の15~20%、視聴時間の0.3~0.8%の増加などを報告した。視聴時間モデルでは演算量FLOPsとパラメータ数をそれぞれ約10%減らしている。依存関係を考慮する過去データ再生ベンチマークでも研究資源の割り当てを評価したが、エージェントが具体的な候補を既に分析・選択する場合、より複雑なスケジューリングによる一貫した効率改善は初期結果では見られなかった。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-24(UTC)
最新改訂
2026-09-24 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Sustaining industrial recommendation research requires using the results of one experiment to decide what to investigate next. We present AgentX-Model, the next generation of AgentX's model research framework, which connects proposal development and model experimentation within sandboxes defined by business inputs and prediction tasks. AgentX-Model adopts a dual-agent architecture comprising a Research Agent and a Model Agent. The Research Agent develops independently reviewed proposals from papers and experimental findings, while the Model Agent conducts multi-round investigations and returns code, measurements, and unresolved questions. Using the returned results, the Research Agent selects a starting implementation and formulates the next research question, allowing subsequent experiments to build on earlier findings. We organize this continuing research around four actions: Reproduce, Follow-up, Composition, and Diagnose. The first three actions drive routine research, while Diagnose acquires the evidence needed to choose a repair, including for issues raised by business feedback and online evaluation, such as prediction bias measured by PCOC. Across the production evaluation, 560 of 636 completed model-changing experiments recorded AUC above their business baselines. As research continued, some experiments recorded AUC above every comparable ancestor in their lineages. The five latest online A/B evaluations across different business settings reported gains including 10-15% in acquisition efficiency, 15-20% in target-segment advertising spend, and 0.3-0.8% in watch time; the watch-time model used approximately 10% fewer FLOPs and parameters. A dependency-aware historical-replay benchmark further evaluates research allocation, with initial results showing no consistent efficiency gain from more complex scheduling when agents already analyze and select concrete candidates.

著者のコメント

Technical report. 37 pages, 11 figures, 13 tables, including appendices

arXiv ID: 2609.30001 / 要約の誤りについて