複数の物理現象が同時に起こる動画の整合性を改善
HiPhy: Hierarchical Alignment for Physically-Plausible Multi-Principle Video Generation
この論文をやさしく読む
ひとことで言うと
一つの動画で複数の物理現象が同時に起きても、それぞれの動きと場面全体が矛盾しにくいように学習する方法です。個々の現象と全体の整合性を別々に評価します。
何に役立つ?
複数の物体や現象が関わる動画生成モデルの学習・評価に役立ちます。5万件のプロンプトと評価用ベンチマークは、単一現象だけでは見えにくい問題を調べるためのものです。
この研究の面白いところ
部分的に自然な動きでも、組み合わせると不自然になる問題に着目しています。局所的な時間変化と場面全体の物理・意味の一致を両方目的に入れています。
どこまで分かった?
要旨は各種ベンチマークでの改善を述べていますが、具体的な改善率や物理量の誤差は示していません。物理的な常識への適合の改善は、厳密な物理シミュレーターとしての正確性を保証するものではありません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
動画生成モデルは高い視覚的忠実度を達成し、汎用的な世界シミュレーターとなる大きな可能性を持つ。しかし、この進展にもかかわらず、物理法則に従う動画の生成にはなお失敗している。同じ動画の中で複数の物理原理が連動しなければならない現実的な状況では、この問題がさらに明確になる。例えば「鍋から蒸気が立ち上る一方で、風船が上方へ浮かぶ」という場面では、浮力と流体力学が整合的に、かつ同時に働く必要がある。ところが既存手法は、動画ごとに単一の原理に注目することが多く、複数の原理の相互作用をほとんど考慮していない。 私たちは、2段階の目的関数を通して動画生成を物理法則に結び付ける強化学習フレームワークHiPhy(Hierarchical Physical Alignment)を提案する。局所的には個々の物理原理の時間的な挙動を守らせ、大域的には場面全体の物理的・意味的な整合性を確保する。複数原理の生成を支援するため、5万件のプロンプトからなるデータセットを構築し、多様な同時発生する物理現象を扱うプロンプトベンチマークMultiPhyBenchを導入する。 実験では、HiPhyは従来手法やベースラインを大幅に上回り、各種ベンチマークで物理的な常識への適合と意味的な整合性を大きく改善した。複数の物理原理が同時に関わり、競合手法の性能が最も急激に低下する場面で、改善が最大だった。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-10-01(UTC)
- 最新改訂
- 2026-10-01 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-10-01 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Video generation models have achieved remarkable visual fidelity and have strong potential to become general-purpose world simulators. Despite this progress, they still fail to generate videos which adhere to laws of physics. The problem becomes even more apparent in realistic settings where multiple physical principles must work together within the same video; for example, "a balloon floating upward while steam rises from a pot" requires buoyancy and fluid dynamics to unfold coherently and simultaneously. Yet existing methods largely ignore multi-principle interactions, focusing on a single principle per video. We propose HiPhy (Hierarchical Physical Alignment), a reinforcement learning framework that grounds video generation in physical laws through a dual-level objective: locally enforcing the temporal dynamics of individual physical principles, and globally ensuring the physical and semantic coherence of the entire scene. To support multi-principle generation, we construct a 50K-prompt dataset and introduce a prompt benchmark MultiPhyBench, spanning a diverse range of co-occurring physical events. Our experiments show that HiPhy significantly outperforms prior methods and baselines, improving physical commonsense and semantic alignment significantly across various benchmarks, with the largest gains on scenes involving multiple concurrent physical principles where competing methods degrade most sharply.
著者のコメント
Project page: https://hiphy-video.github.io/
arXiv ID: 2610.02197 / 要約の誤りについて