単一の車窓映像から終わりのない風景と音を生成する作品
Passing: An Endless Journey through Reconstructed Spacetime with AI-Generated Sound
この論文をやさしく読む
ひとことで言うと
一本の車窓映像を時間と空間で組み替え、鑑賞者の在席に応じて映像と生成音を変える作品。
何に役立つ?
映像の再構成、生成音、観客の反応を組み合わせるインタラクティブ作品の制作・研究例になる。
この研究の面白いところ
AIを正しい音の復元器ではなく、変形した映像を音として解釈する役割に置いている。
どこまで分かった?
論文要旨は作品の仕組みと創作上の考えを述べる。鑑賞者への効果を定量評価した結果は記載されていない。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
本論文は、一本の連続したモノレールの車窓映像を時空間の立体として再構成し、終わりのない旅を生成する対話型の映像・音響作品Passingを紹介する。映像を時間順に再生する代わりに、その空間・時間構造を非線形の経路に沿って再サンプリングする。これにより、奥行き、速度、時間の順序が不安定に変わりながら、風景が連続して流れる。カメラによる鑑賞者の在席検出システムは、鑑賞者が鑑賞区域にいるかを推定し、その状態を使って描画される映像列の切替えに影響を与える。得られた映像はリアルタイムの動画から音声を合成するSpecMaskFoleyへ送られ、組み替えた映像に同期する音環境が生成される。 このモデルは客観的に正しい音声を復元するためのものではなく、通常の空間・時間の前提が崩れた世界について、あり得る聴覚的解釈を提案する仮説的な聞き手として働く。作品では、時空間の再構成規則を定める作家、現れた視覚の流れを音として解釈するAIモデル、そして身体的な存在によって映像と音の経路に影響を与える鑑賞者の間に、創作上の主体性を分ける。この構成を通じて、人の意図、機械の推論、鑑賞者の解釈の間で、作者性と聴く行為がどう形づくられるかを探る。作品ページも公開されている。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-23(UTC)
- 最新改訂
- 2026-09-23 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-23 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
This paper introduces Passing, an interactive audiovisual installation that generates an endless journey from a single continuous monorail-window recording by reconstructing it as a spatiotemporal volume. Rather than replaying the footage linearly, the work resamples its spatial and temporal structure along nonlinear trajectories, producing a continuously passing landscape whose depth, speed, and temporal order become unstable. A camera-based viewer-presence detection system estimates whether a viewer is present in the viewing zone and uses this presence state to influence transitions among rendered video sequences. The resulting video stream is fed into SpecMaskFoley, a real-time video-to-audio synthesis model that generates a synchronized soundscape for the reconfigured image. The model is not used to reconstruct an objectively correct soundtrack, but functions as a speculative listener, proposing a possible auditory interpretation of a world whose conventional spatial and temporal premises have been disrupted. Passing distributes creative agency across the artist, who defines the rules of spacetime reconstruction; the AI model, which interprets the emergent visual flow as sound; and the audience, whose embodied presence influences the audiovisual trajectory. Through this structure, the work investigates how authorship and listening may be negotiated among human intention, machine inference, and audience interpretation. Artwork page: https://ryufurusawa.com/passing
著者のコメント
Accepted to the NeurIPS 2026 Creative AI Track
arXiv ID: 2609.27489 / 要約の誤りについて