arXiv論文メモ
新着一覧
cs.CV / cs.AI · 査読状況未確認

評価基準を生成して動画の報酬モデルを安定させる手法

RewardVerse: Rubric-Guided Policy Optimization for Video Reward Modeling

Zhenchen Tang, Yang Li, Songlin Yang, Bo Peng, Xiaotong Zhao, Shuai Li, Haotian Fan, Alan Zhao, Jing Dong

この論文をやさしく読む

ひとことで言うと

動画の良し悪しを一つの点数で評価する前に、評価基準を明示して採点を安定させる報酬モデルの研究である。

何に役立つ?

動画生成の強化学習で、採点尺度の揺れを抑え、人の評価に沿う報酬を設計する際に役立つ可能性がある。実際の生成品質への効果は要旨だけでは数値で示されていない。

この研究の面白いところ

評価基準の生成器と採点器を二段階で学習し、問い合わせに応じて評価項目を変えつつ、人間の評価との整合を保つ。

どこまで分かった?

要旨はベンチマークと外部データセットでの評価を述べるが、具体的な改善数値や実際の強化学習後の動画品質は記載していない。

v2のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

動画生成モデルの最適化には強化学習が重要であり、頑健な報酬モデルがその基盤となる。しかし既存の動画報酬モデルは、複雑で主観的な動画品質を明示的な評価基準なしに一つの数値へ直接写すため、スコアが不安定になることが多い。その結果、プロンプトごとに採点尺度がつぶれたり移動したりするスカラーのドリフトが起き、強化学習の報酬として信頼しにくくなる。人による専門的な注釈設計を参考に、この問題へ対処するRewardVerseを提案する。これは評価の問い合わせと採点器の間に、動的な評価基準を中間表現として置く動画報酬の枠組みである。無制約に直接採点する代わりに、まず明示的な評価基準を生成し、それに沿って採点する。意味上の安定した基準点を与え、スカラーのドリフトを抑える。この協調的な処理過程を効率よく最適化するため、二段階の学習アルゴリズムRubric-Guided Policy Optimization(RGPO)を提案する。まず自ら発展する初期の評価基準を使って採点器を準備し、次に、問い合わせに適応する評価基準を生成するよう生成器を共同で最適化しながら、採点器を人間の評価に継続して合わせる。16次元のEvalVerseベンチマークと外部データセットでの広範な実験は、RewardVerseがスカラーのドリフトを抑え、単体評価と一対比較の双方で最先端の性能を達成し、動画生成の強化学習に向けた頑健で解釈可能な報酬信号を与えることを示す。

v2の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-19(UTC)
最新改訂
2026-09-23 · v2
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Reinforcement learning (RL) is vital for optimizing video generation models, with a robust reward model (RM) serving as the cornerstone. However, existing video reward models often produce unstable scalar scores because they directly map complex, subjective video quality into a single score without explicit evaluation criteria. This leads to scalar drift, where the scoring scale collapses or shifts across different prompts, making the reward unreliable for RL. Drawing inspiration from professional human annotation engineering, we address this problem with RewardVerse, a rubric-based video reward framework that introduces a dynamic rubric as an intermediate representation between the evaluation query and the scorer. Instead of unconstrained direct scoring, RewardVerse first generates explicit evaluation criteria and then performs rubric-guided scoring, providing a stable semantic anchor that mitigates scalar drift. To efficiently optimize this collaborative pipeline, we propose Rubric-Guided Policy Optimization (RGPO), a two-stage training algorithm. RGPO first warms up the scorer using self-evolving seed rubrics and then jointly optimizes the rubric generator to produce query-adaptive evaluation criteria while continuously aligning the scorer with human ratings. Extensive experiments on the 16-dimensional EvalVerse benchmark and external datasets demonstrate that RewardVerse mitigates scalar drift, achieves state-of-the-art performance on both pointwise and pairwise evaluation, and provides a robust and interpretable reward signal for RL in video generation.

arXiv ID: 2609.22947 / 要約の誤りについて