一人称と外部視点が変わる動画の継続学習を評価する
CE$^4$L: Continual Ego, Exo, and Ego-Exo Learning
この論文をやさしく読む
ひとことで言うと
見る位置と取り組む課題が同時に変わる動画を、順番に学習するモデルの評価環境です。
何に役立つ?
一人称映像と外からの映像を扱う継続学習で、従来の単一視点評価では見えにくい弱点を調べるのに役立ちます。
この研究の面白いところ
課題ごとの小さなアダプターを保存し、統計量から作る部分空間への距離で利用するアダプターを選びます。
どこまで分かった?
学習不要なのは経路選択の部分で、モデル全体が学習不要という意味ではありません。要旨には性能の具体値や実機での運用結果はありません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
身体を持つエージェントの知覚は動画に基づき、一人称視点、外部視点、またはその両方を使う多視点であることが多く、課題と視点が同時に変化する本質的に継続的なものである。しかし継続学習(CL)では、外部視点だけの認識課題が依然として中心であり、こうした現実の複合的な変化のもとでの振る舞いは分かりにくい。 本研究では、一人称・外部視点・両者を組み合わせた継続学習を扱うCE⁴Lを導入する。これは、視点をまたぐ参照に基づく技能評価、時間的動作分割、視点間の対応付け、行動予測と計画という4つの代表的課題を含む、統一的な多視点CLベンチマークである。CE⁴Lは、視点間対応、視点に依存した非同期性、異種の意味的目標など、従来のCLベンチマークではほとんど扱われなかった課題を明らかにする。 これに対し、動画向け増分学習の部分空間経路選択型タスクアダプターVISTAを提案する。これは課題固有の更新を軽量アダプターに保存するパラメーター効率のよい基準手法であり、2次統計量から推定した課題固有の白色化部分空間への残差距離に基づき、経路選択のための学習を行わずに振り分ける。広範な実験により、代表的なCL手法の有効性がCE⁴Lの設定間で大きく異なる一方、VISTAは一貫して競争力があり、全体として最先端の性能を達成することを示す。ベンチマークと手法のソースコードはhttps://github.com/AnAppleCore/CE4Lで公開している。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-20(UTC)
- 最新改訂
- 2026-09-20 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-20 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Perception for embodied agents is video-based, often multi-view (ego, exo, or both), and inherently continual, with simultaneous task and viewpoint shifts. Yet continual learning (CL) remains dominated by exo-only recognition tasks, obscuring behavior under these real-world coupled shifts. We introduce Continual Ego, E}xo, and Ego-Exo Learning (CE$^4$L), a unified multi-view CL benchmark spanning four representative tasks: cross-view referenced skill assessment, temporal action segmentation, cross-view association, and action anticipation & planning. CE$^4$L highlights challenges largely absent in prior CL benchmarks, including cross-view correspondence, view-dependent asynchrony, and heterogeneous semantic objectives. To this end, we propose Video Incremental Subspace-routed Task Adapters (VISTA), a parameter-efficient baseline method that stores task-specific updates in lightweight adapters and performs training-free routing via residual distance to task-specific whitened subspaces estimated from second-order statistics. Extensive experiments demonstrate the significantly varied efficacy of representative CL methods across CE$^4$L settings, while VISTA is consistently competitive and achieves state-of-the-art overall performance. Our source code for benchmarks and methods is available at https://github.com/AnAppleCore/CE4L .
著者のコメント
23 pages. Accepted by ICML 2026
arXiv ID: 2609.23492 / 要約の誤りについて