arXiv論文メモ
新着一覧
cs.AI · 査読状況未確認

予算付き配達員報酬計画を依頼ごとに作る枠組み

A General Framework for Budgeted Threshold Incentives on Request

Zhuolin Wu, Chengrui Zhu, Wenhua Nie, Kenny Ye Liang, Junming Lin, Haiyang Li, Zhilin Li, Wenjia Geng, Zeyu Wu, Yinan Wu, Jinghua Hao, Renqing He

この論文をやさしく読む

ひとことで言うと

配達員向けの段階別報酬を、期間や予算が変わる依頼ごとに速く計画する方法。

何に役立つ?

無作為化試験を長期間待てないとき、限られた予備試験と既存履歴を使って報酬計画を比較するのに役立つ。

この研究の面白いところ

4段階の誤差を理論的に分け、3,000人のデータや制御した応答法則で速度と後悔値を評価した。

どこまで分かった?

保証は固定された計画候補と段階ごとの誤差が与えられる条件下。実際の運用成果がすべての都市や条件で再現するとは示していない。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

オンデマンド配送プラットフォームは、似た履歴の配達員による最近の完了件数から段階を定めた報酬活動で配達員に支払う。運営者は祝日や悪天候などに合わせて期間、対象配達員、支払規則、予算が異なる計画を求めるが、無作為化試験は少なく、収集に数か月を要する。著者らは、条件付き予測、対象集団の縮約、軌跡の積分、予算配分という4段階を、交換可能な7つのモジュールで構成する依頼駆動型の枠組みを提示する。モジュール間では条件付き軌跡の確率法則を交換し、その報酬確率と報酬で印を付けたモーメントから、どの活動規則でも支払額と増分効果を求められる。応答補正では、豊富な報酬なしの履歴から得た軌跡に重みを付け直し、短い予備試験のモーメントへ合わせる。固定された計画候補と各段階の誤差が与えられた場合、全体の価値損失は4段階の誤差項の和で抑えられると証明する。また、どの段階についても、それを省くと他段階で取り除けない誤差の下限が残る例を示す。3,000人の配達員の45週間の開始時点を対象に、1週間内の全127期間に対して、同一のシナリオで11.04倍速く回答し、配分による価値損失は最大0.92%だった。新たに制御した24種類の応答法則では、1週間の予備試験を使う応答補正が、名目上同じ無作為化配達員週数の試験に比べて後悔値を51.2%下げた。4週間の予備試験と厳密な総和では、18週間の試験との差が+0.007以内だった。期間、集団、規則、拘束的な予算が依頼ごとに変わる登録済みの研究では、この枠組みの後悔値は同じ名目配達員週数の試験と、同じ予備試験データの用量補間より低かった。一度だけの準備を再利用すると、60件の依頼に同じ回答をそれぞれ14.1倍、2.70倍速く返せた。枠組み自身の用量曲線で当てはめた9種類の報酬案の試験に比べ、1週間時点の後悔値は0.055低かった。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-24(UTC)
最新改訂
2026-09-24 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

On-demand delivery platforms pay riders through incentive activities whose tiers are set from recent completions of riders with a similar history. Operators request such plans for changing periods, rider populations, payment rules and budgets, often for holidays or bad weather, where randomized trials are scarce and take months to collect. We present a request-driven framework that composes four stages (conditional prediction, population reduction, trajectory integration and budget allocation) through seven replaceable modules that exchange conditional trajectory laws, whose award probabilities and award-marked moments give payment and uplift for any activity rule. A response-correction step reweights trajectories from abundant no-offer history to match the moments of a short pilot. We prove that, on a fixed plan menu and given the stage errors, the end-to-end value loss is bounded by the sum of four stage terms, and that for every stage there are instances on which omitting it leaves an error floor the others cannot remove. On 3,000 riders over 45 weekly origins, all 127 windows of a week are answered 11.04x faster with identical scenarios and at most 0.92% value lost by the allocation. On 24 new controlled response laws, the response correction with a one-week pilot lowers regret by 51.2% relative to a trial with the same nominal randomized rider-weeks, and a four-week pilot with exact summation comes within +0.007 of an 18-week trial. In registered studies where windows, populations, rules and binding budgets change from request to request, the framework's regret is below that of a trial with the same nominal rider-weeks and below dose interpolation of the same pilot data, and reusing its one-off preparation answers 60 requests 14.1x and 2.70x faster with identical answers. Against a nine-offer trial fitted with the framework's own dose curve, one-week regret is 0.055 lower.

著者のコメント

42 pages

arXiv ID: 2609.29724 / 要約の誤りについて