長い動画で全点を追跡する3D表現TrackEverything
TrackEverything: Long Horizon Dense Tracking via De-Duplicating 3D Scene Representations
この論文をやさしく読む
ひとことで言うと
動画中に見える多数の点を長時間追えるよう、同じ3D表面の重複した観測を統合する追跡方法。
何に役立つ?
長い動画での3D点追跡の評価・開発に役立つ。要旨ではTAPVid-3Dでの比較とGPUメモリー条件を示す。
この研究の面白いところ
計算量をフレーム数ではなく、固有のシーン形状に結び付ける。40 GBのGPUメモリーで1,000フレーム超の全可視点を追えると報告する。
どこまで分かった?
短い映像での20%超というAPD差は公開されている全フレーム密3D追跡器との比較。長い系列では疎な追跡器と競争力があるという記述で、同じ点数を追跡した比較ではない。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
従来の点追跡モデルには、少数の指定点を長時間追うか、短い映像で全点を追うかという根本的な両立の難しさがある。本研究は、動画を世界座標における持続的な3Dシーンの軌跡として表すことで、この制約を乗り越える3D点追跡器TrackEverythingを導入する。動画は背後にある3D世界の2D投影だという考えに基づき、モデルの複雑さを動画の長さから切り離し、異なる実物のシーン形状の量に応じて増えるようにする。方法には三つの工夫がある。第一に、スライド窓の境界でボクセル化による重複除去を行い、同じ場所にある軌跡を統合して、同じ表面を繰り返し観測しても情報が冗長に蓄積しないようにする。第二に、各点の到達位置と静止・移動の分類を予測する端点洗練器と、移動する点についてだけ密な軌跡を復号する軽量な軌跡洗練器に処理を分ける。第三に、メモリーを大量に使う4D相関ボリュームを、シーン点群からの効率的な特徴サンプリングに置き換える3D WAFTを提案する。著者らの知る限り、40 GBのGPUメモリー内で1,000フレームを超える動画の可視点をすべて追跡できる最初の3D追跡器である。TAPVid-3Dの短い映像では、公開されている全フレーム密3D追跡器のすべてをAPDで20%超上回った。長い系列では、はるかに多くの点を追跡しながら、最先端の疎な点追跡器と競争力のある結果を保った。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-24(UTC)
- 最新改訂
- 2026-09-24 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-24 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Existing point tracking models face a fundamental tradeoff: they can either track a sparse set of query points over long horizons, or track all points across only short clips. We introduce TrackEverything, a 3D point tracker that breaks this trade-off by representing videos as persistent 3D scene tracks in world coordinates. Grounded in the insight that videos are 2D projections of an underlying 3D world, TrackEverything decouples model complexity from video duration, allowing it to scale with unique physical scene geometry instead. Our approach introduces three key innovations. First, we employ a voxelization-based de-duplication mechanism at sliding-window boundaries to merge co-located tracks, preventing repeated observations of the same surface from redundantly accumulating. Second, we decompose tracking into an endpoint refiner that predicts each point's destination and static-versus-dynamic classification, followed by a lightweight trajectory refiner that decodes dense trajectories exclusively for dynamic points. Third, we propose 3D WAFT, replacing memory-prohibitive 4D correlation volumes with efficient feature sampling in the scene cloud. To the best of our knowledge, TrackEverything is the first 3D tracker capable of tracking all visible points across videos exceeding 1000 frames within 40 GB of GPU memory. On TAPVid-3D, TrackEverything outperforms all open-source all-frame dense 3D trackers by more than 20% APD on short clips, while remaining competitive with state-of-the-art sparse trackers on long sequences, despite tracking far more points.
arXiv ID: 2609.30222 / 要約の誤りについて