反復動作の構造を保つ動画合成で回数計測を学習
TReViS: Temporal Repetition Structure Aware Video Synthesis for Self-supervised Repetitive Action Counting
この論文をやさしく読む
ひとことで言うと
動画の中で動作が何回繰り返されたかを数えるモデルを、人が反復回数や区間を教えずに学習させます。元動画の周期構造を推定し、学習に使う新しい動画と擬似ラベルを作ります。
何に役立つ?
反復動作の注釈を大量に用意する負担を減らす方法として役立ちます。考えられる用途は、運動や繰り返し作業を含む動画の回数計測です。
この研究の面白いところ
単純に映像を繰り返すだけでなく、元の反復構造を維持しながら時間的な変動を加えます。合成データを使い、既存の回数計測モデルをゼロから学習できる点が特徴です。
どこまで分かった?
教師あり手法のすべてを上回ったという結果ではなく、いくつかと競争力があると報告しています。要旨には具体的な誤差値や、対象データセット以外での性能は記載されていません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
完全教師ありの反復動作回数計測(RAC)は高い性能を達成しているが、時間方向に密な注釈が必要であり、その作成は高コストで規模を拡大しにくい。本研究では、反復に関するラベルを一切用いずにRACモデルを学習できる、自己教師あり動画合成フレームワークTReViSを提案する。TReViSは、時間的自己類似度行列を用いて、ラベルのない動画に内在する時間的な反復構造を推定する。そして周期の統計量を推定し、現実的な反復パターンを保持しつつ、制御された時間的変動を導入した新たな学習系列を合成する。 これらの合成動画に擬似ラベルを組み合わせ、既存のRAC構成をゼロから学習させる。複数のデータセットとバックボーンにわたり、TReViSは従来の自己教師あり手法を一貫して上回り、完全にラベルを用いないまま、いくつかの教師あり比較手法と競争力のある性能を達成する。これにより、反復構造を考慮した動画合成が、ラベルなしのRACに有効であることを示す。ソースコードはhttps://github.com/yfqi/TReViSで公開されている。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-21(UTC)
- 最新改訂
- 2026-09-21 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-21 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Fully supervised repetitive action counting (RAC) has achieved strong performance, but requires dense temporal annotations that are costly and difficult to scale. We propose TReViS, a self-supervised video synthesis framework that enables training RAC models without any repetition labels. TReViS estimates the underlying temporal repetition structure of an unlabeled video via a Temporal Self-Similarity Matrix, infers its cycle statistics, and synthesizes new training sequences that preserve realistic repetition patterns while introducing controlled temporal variability. These synthesized videos are paired with pseudo-labels and used to train existing RAC architectures from scratch. Across multiple datasets and backbones, TReViS consistently outperforms prior self-supervised methods and achieves performance competitive with several supervised baselines, while remaining fully label-free, demonstrating the effectiveness of structure-aware video synthesis for label-free RAC. The source code is available at https://github.com/yfqi/TReViS.
著者のコメント
Accepted for publication in Image and Vision Computing (Elsevier). This is the author-accepted manuscript and not the final published version of record. The DOI and link to the published version will be added when available
arXiv ID: 2609.24367 / 要約の誤りについて