映画の音声ガイドで内容・時刻・言い方を同時に決める
What, When, and How: Audio Description as Constrained Global Optimization
この論文をやさしく読む
ひとことで言うと
映画の音声ガイドについて、説明する内容と台詞の合間の時刻、短い言い方をまとめて決める方法です。
何に役立つ?
視覚障害者向けの音声ガイドを自動作成する際、台詞と重ならず物語に必要な情報を選ぶ設計に役立ちます。
この研究の面白いところ
言語モデルで候補を作り、混合整数線形計画で場面全体の説明時刻を決める構成です。
どこまで分かった?
REFRAMEDでの時刻・物語の指標は改善しましたが、専門の制作者との間にはなお大きな差があります。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
音声ガイド(AD)は、台詞の合間に映像情報を語ることで、視覚障害のある人が映画を利用できるようにする。従来の自動ADシステムは、何を説明するかと説明する時刻がすでに与えられていると仮定し、局所的な映像から文章への変換として扱うことが多い。しかし現実のADでは、物語に重要な映像情報は何か、台詞を妨げずにいつ話せるか、限られた時間に収まるようどう表現するかを同時に決める必要がある。本研究は、これら3つの判断に対する制約付き最適化問題としてAD生成を定式化する。ハイブリッドシステムは、大規模言語モデルを使って映像要素を提案して映像に対応付け、物語上の重要度を推定し、短くした説明文を生成する。続いて混合整数線形計画によって、時間の制約の下で場面全体にわたる説明の選択と配置を同時に行う。現実的な映画ADのベンチマークREFRAMEDで評価したところ、プロンプトを与えたLLMより、何をいつ説明するかの判断が良く、物語に関する質問応答と時刻への対応を評価する指標で新たな最高水準となった。要素を取り除いた評価では、時間制約を明示することが配置の改善を生み、重要度推定が物語に役立つ内容をどれだけ残せるかを左右することが分かった。改善はnグラムの一致よりも時刻や物語に関する指標に集中したが、専門の音声ガイド制作者との間にはなお大きな差がある。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-24(UTC)
- 最新改訂
- 2026-09-24 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-24 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Audio Description (AD) makes movies accessible to blind and visually impaired audiences by narrating visual information in gaps between dialogue. Existing automatic AD systems largely treat generation as a local video-to-text problem, assuming that the content to describe and its temporal location are already provided. Realistic AD instead requires coupled decisions about what visual information is narratively important, when it can be spoken without interfering with dialogue, and how it should be formulated to fit within the available time. We formalize AD generation as a constrained optimization problem over these three decisions. Our hybrid system uses large language models to propose and ground visual elements, estimate their salience to the narrative, and generate compressed realizations. A mixed-integer linear program then jointly selects and schedules descriptions across a scene subject to temporal constraints. When evaluated on REFRAMED, a benchmark for realistic AD of movies, our approach makes better decisions than prompted LLMs about what to describe and when to describe it, establishing a new SOTA on narrative QA and temporally grounded metrics. Ablations show that explicit temporal constraints drive gains in placement, while salience estimation controls how much narratively useful content is retained. Improvements are concentrated on temporal and narrative measures rather than n-gram overlap, although a significant gap to professional describers remains.
arXiv ID: 2609.30121 / 要約の誤りについて