生成フローモデルをシミュレーションなしで報酬に合わせるWTF
WTF?! Simulation-Free Reinforcement Learning with Wasserstein-Tilted Flow Maps
この論文をやさしく読む
ひとことで言うと
生成画像などのフローモデルを、個々の出力を高報酬へ動かす形で効率よく追加学習する方法。
何に役立つ?
画像生成モデルを少ない推論ステップで報酬に合わせたい場合の、考えられる学習手法。
この研究の面白いところ
分布の重み付けではなく最適輸送で出力を動かし、決定的な最適制御との同値性を使う。
どこまで分かった?
計算量最大280分の1はImageNet-256と文章画像生成で比較した条件の結果。すべての生成課題への一般化は要旨にない。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
報酬に基づく追加学習は、事前学習済みのフロー型生成モデルを更新し、生成結果の後段の報酬を高めることを目指す。従来は、KL正則化付きの報酬最大化問題の解となる、報酬に応じて重みを変えた分布から標本を取る問題として定式化することが多い。本研究では、事前学習したドリフトから直接作る最適輸送の正則化を導入する。これにより、元の分布の重みを変えるのではなく、個々の標本をより高い報酬へ移動させる目的関数になる。この問題は、フローに対する決定的な最適制御問題と同値であることを示す。事前学習済みのフローマップがあれば、この同値性から、生成フローを追加学習するシミュレーション不要の強化学習アルゴリズムが得られる。 この枠組みをWasserstein-Tilted Flow Maps(WTF)と呼び、フローマップに元から対応した端から端までの追加学習手順として提示する。できあがったフローマップは、後から蒸留しなくても、少ない推論ステップで報酬に沿う高い性能を保つ。ImageNet-256と文章からの画像生成の実験では、比較方法と同程度かそれ以上の多様性を保ちつつ、より高い報酬を達成し、学習計算量を最大280分の1に減らした。さらに、フローマップのような高速サンプラーは効率的な追加学習の重要な基盤であり、主流のKL正則化による定式化も、再検討すべき複数の選択肢の一つだと論じる。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-22(UTC)
- 最新改訂
- 2026-09-22 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-22 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Reward fine-tuning aims to update a pre-trained flow-based generative model to improve the downstream reward of its generated samples. Existing methods typically formulate this problem as sampling from a reward-tilted distribution, the solution to a KL-regularized reward-maximization problem. Here, we introduce an optimal transport regularizer built directly from the pre-trained drift. Unlike KL reward tilting, the resulting objective transports individual samples toward higher reward rather than reweighting the base distribution. We show that the resulting problem is equivalent to a deterministic optimal control problem on the flow. Given a pre-trained flow map, this equivalence yields a simulation-free reinforcement learning algorithm for fine-tuning generative flows. We call the resulting framework Wasserstein-Tilted Flow Maps (WTF), the first end-to-end fine-tuning recipe native to flow maps. The output is a fine-tuned flow map that retains strong reward-aligned performance at few-step inference budgets without post-hoc distillation. Experiments on ImageNet-256 and text-to-image show that WTF achieves higher reward with comparable or higher diversity than baselines, while requiring up to $280\times$ less training compute. More broadly, we argue that accelerated samplers such as flow maps are essential infrastructure for efficient post-training, and that the dominant KL-regularized formulation is only one of many choices worth revisiting.
arXiv ID: 2609.27033 / 要約の誤りについて