arXiv論文メモ
新着一覧
cs.CV · 査読状況未確認

地面の格子と円柱で人物とカメラの動きを指定する動画生成

CoaG: Cylinders on a Grid for Coarse 3D Layout Control in Video Generation

Zhangsihao Yang, Mengyi Shan

この論文をやさしく読む

ひとことで言うと

地面の格子と人物ごとの円柱を動かして、生成動画内の人物配置とカメラ移動を指定する方法。

何に役立つ?

考えられる用途は、動画制作時に人物の配置やカメラの経路を簡単な図で指定すること。要旨では、合成データで学習したモデルが評価用動画で指定した配置や複数のカメラ移動に従うかを確かめている。

この研究の面白いところ

対応する実データがないため、2000件の説明文から動画を作り、そこから人物と地面の幾何情報を逆に取り出して学習組を作った。実写映像や手作業のラベルを使っていない。

どこまで分かった?

評価は合成した組から取り分けた動画で行われた。カメラを遠ざける動きへの追従は弱かった。実写データでの性能や、要旨にない複雑な場面での性能は示されていない。

v2のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

生成動画で人物が立つ位置とカメラの動きを制御するには、どれほど少ない幾何情報を描けばよいのか。本研究の答えは、地面の平面と、人物一人につき一本の円柱である。利用者は地面に格子を描き、各人を立たせたい位置に円柱を置き、81フレームにわたって円柱とカメラを動かす。モデルは、人物が円柱の位置を占め、円柱に合わせて移動し、指定したカメラから見える写実的な動画を生成する。人物の外見は文章による指示と背景の参照画像から、配置と動きは幾何情報から決まる。 このような制御信号と動画を対応付けたデータセットがないため、著者らは自ら組を作った。自動処理系が組合せ的な種から2000件の説明文を作り、それぞれをテキストから動画を生成するモデルで映像化した。さらに人物追跡、背景の補完、エージェントを使った地面マスク処理、多視点のフィードフォワード再構成、平面当てはめによって、各動画から幾何情報を復元した。実写映像も手作業のラベルも使っていない。こうして得た1935組でWan2.2-Fun-ControlのLoRAを学習させたところ、評価用に取り分けた動画で描いた配置とカメラ経路に従った。生成人物は円柱の数、順序、位置、高さに対応し、文章で人物の属性が変わり、参照画像で背景が変わった。カメラを近づける、周囲を回る、左右に振る、上下させる経路には従ったが、遠ざける動きへの追従は弱かった。

v2の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-21(UTC)
最新改訂
2026-09-23 · v2
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

We ask how little geometry a person has to draw to control both where people stand and where the camera moves in a generated video. Our answer is a ground plane and one cylinder per person. A user draws a grid on the ground, places one cylinder where each person should stand, moves the cylinders and the camera over 81 frames, and the model renders a photoreal video in which the people occupy the cylinders' positions, move as the cylinders move, and are seen from the drawn camera. Appearance comes from a text prompt and a background reference image; layout and motion come from the geometry. Because no dataset pairs such a signal with video, we build the pairs ourselves: an automatic engine writes 2000 captions from a combinatorial seed, generates a clip for each with a text-to-video model, and lifts every clip back to its geometry with person tracking, background inpainting, an agentic ground-mask loop, feed-forward multi-view reconstruction and a plane fit, with no real footage and no manual labels. A LoRA on Wan2.2-Fun-Control trained on 1935 such tuples follows drawn layouts and camera paths on hold-out clips: the generated people match the cylinders' count, order, position and height, the text changes who they are, the reference image changes where they are, and dolly-in, orbit, pan and crane paths are followed, dolly-out only weakly.

著者のコメント

Project page with videos: https://zshyang.github.io/CoaG/

arXiv ID: 2609.24208 / 要約の誤りについて