arXiv論文メモ
新着一覧
cs.CV · 査読状況未確認

手持ちの動画に登場する人物や場所を保って補助映像を生成する

Memory-Guided B-Roll Generation from User Video Collections

Cusuh Ham, Fabian Caba Heilbron, Josef Sivic, Bryan Russell

この論文をやさしく読む

ひとことで言うと

手持ちの映像から人物や場所の情報を整理し、足りない補助ショットを生成するシステムです。

何に役立つ?

動画編集で、撮影素材の雰囲気を引き継いだ補助映像を作る用途を想定しています。ユーザー評価では指示への適合性に利点がありました。

この研究の面白いところ

生成だけに任せず、素材から参照フレームを検索し、ショット間の一貫性を繰り返し確認します。

どこまで分かった?

生成方式が全指標で優位だったわけではありません。撮影素材の検索だけで作った系列は、視覚的整合性で58.2%の比較において好まれました。調査人数などの詳細は要旨にありません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

ユーザーの動画コレクションに根拠を置くBロール系列生成の手法を紹介する。ユーザーの動画群、自然言語による指示、目標の長さを入力とし、主要映像であるAロールを補完しつつ、コレクション内の人物、場所、物体、スタイルを保持した複数ショットの系列を生成することが目的である。この課題は、何時間もの撮影映像から、系列の各ショットの生成を導く視覚的な根拠を選ばなければならないため難しい。 本研究ではMemComposerでこの課題に対処する。これは、生の映像を人物、場所、物体、スタイルの視覚的な参照を持つ構造化メモリへ変換し、それを使ってコレクションに基づくBロール系列を計画、検索、生成する三段階のシステムである。第一に、一度だけ行うオフライン処理で、生の動画から実体を中心としたメモリを構築する。第二に、そのメモリとユーザーの指示から根拠のある系列を計画し、各ショットの条件付けに用いるフレームを検索する。第三に、生成と評価を反復し、同一性、場所、系列全体の一貫性を確保する。 ユーザーの選好調査により、指示への適合性と動画コレクションとの視覚的な整合性の二つの軸でMemComposerを評価した。コレクションに基づかないテキストから動画への計画手法との比較では、指示への適合性の60.0%、視覚的整合性の92.8%の比較でMemComposerが選ばれた。これはコレクションのメモリと参照検索による根拠付けの利点を示す。撮影済み映像を検索して組み合わせるだけの系列との比較では、指示への適合性の94.5%でMemComposerが選ばれ、不足するショットを生成する価値が示された。一方、視覚的整合性については、58.2%の比較で検索のみの系列が好まれた。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-10-01(UTC)
最新改訂
2026-10-01 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

We introduce an approach for collection-grounded B-roll sequence generation. Given a user's video collection, a directive given in natural language, and a target duration, the goal is to produce a multi-shot sequence that complements the user's primary footage (A-roll) while preserving the collection's characters, settings, objects, and style. This task is challenging as one must choose the visual evidence from hours of captured footage that should guide the generation of each shot in the sequence. We address this challenge with MemComposer, a three-stage system that turns raw footage into a structured memory with visual references (characters, settings, objects, and style) and uses it to plan, retrieve, and generate grounded B-roll sequences. First, in a one-time offline stage, MemComposer constructs an entity-centric memory from raw video. Second, it uses the memory and user directive to plan a grounded sequence and retrieve conditioning frames for each shot. Third, it iteratively generates and critiques the sequence to enforce identity, setting, and sequence-level consistency. We evaluate MemComposer in a user preference study along two dimensions: prompt adherence and visual alignment to the user's collection. Against an ungrounded text-to-video planner, MemComposer wins 60.0\% of prompt-adherence and 92.8\% of visual-alignment comparisons, showing the grounding benefit of collection memory and reference retrieval. Against retrieval-only sequences assembled from captured footage, MemComposer wins 94.5\% of prompt-adherence comparisons, showing the value of generating missing shots, while retrieval-only sequences are preferred for visual alignment in 58.2\% of comparisons.

著者のコメント

Project page at https://cusuh.github.io/MemComposer

arXiv ID: 2610.01884 / 要約の誤りについて