追加学習なしでクライミングのホールド使用を検出
Training-Free Hold-Usage Detection in Sport Climbing with Foundation Pose Models
この論文をやさしく読む
ひとことで言うと
単眼動画と既存の姿勢モデルだけで、クライマーの指先やつま先が使うホールドを検出した。
何に役立つ?
装着機器を使わない登攀の採点や動作分析に向けた方法として検討できる。
この研究の面白いところ
複雑な特徴判定や身体部位の分割を加えるより、指先・つま先の点と簡単な時間規則だけの方が良かった。
どこまで分かった?
評価は22動画、10人、2ルートのデータ。イベントF1は評価分割と時間的な重なりの条件で変わる。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
登山者がどのホールドをいつ使うかの検出は、スポーツクライミングの自動採点、動作分析、支援システムの基礎となる。既存法は専用モデルを学習するか、二次元姿勢推定を転用していた。しかし後者の手の点は手首、足の点は足首にあり、実際にホールドに触れる指先やつま先からずれている。手は全フレームのおよそ半分で隠れている。本研究は、追加学習せず固定した既存の姿勢基盤モデルで十分であることを示す。Sapiensの指先とつま先のキーポイントについて、注釈付きホールドとのフレームごとの近接判定、同じ手足で複数を同時に使わない規則、短時間の持続条件を組み合わせ、クライミング専用学習なしで使用を検出した。 22本の動画、10人の選手、2ルートからなるThe Way Upデータセットでは、保留した分割でイベントF1が90.2%、参加者を一人ずつ除く交差検証で89.8%だった。22本全体で時間区間に少しでも重なりがあればよい条件では79.9%だった。足のホールドが最も良く、全体でF1 89.8%、保留分割で96.6%だった。同じ評価手順で再現したYOLOv8-poseとViTPoseの処理系列を、すべての時間閾値で上回り、時刻の厳密さを増すほど差が広がった。構成要素を除く実験では、基盤モデルの密な特徴の変化による判定と身体部位の分割は、直感に反して性能を下げた。自動予測から求めた標準的な指導用統計も正解データに近く、登攀時間のPearson相関は1.00、ペースでは0.94だった。通常の単眼動画から、装着機器なしで成績指標を得られることを示す。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-24(UTC)
- 最新改訂
- 2026-09-24 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-24 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Detecting which holds a climber uses, and when, underpins automated scoring, movement analysis, and assistive systems for sport climbing. Existing approaches train task-specific models or repurpose 2D pose estimators whose hand keypoint sits at the wrist and foot keypoint at the ankle i.e. offset from the fingertips and toes that actually contact the holds, and whose hands are occluded in roughly half of all frames. We show that a frozen, off-the-shelf pose foundation model is sufficient: using the fingertip and toe keypoints of Sapiens, a per-frame proximity test against the annotated holds, per-limb mutual exclusion, and a short temporal-persistence rule, we detect hold usage without any climbing-specific training. On the The Way Up dataset (22 videos, 10 athletes, two routes), our method reaches an event F_1 of 90.2% on a held-out split (89.8% under leave-one-participant-out cross-validation) and 79.9% over all 22 videos at any temporal overlap, and performs best on footholds (F_1,89.8% overall, 96.6% held-out). Under an identical protocol it exceeds our reproductions of the YOLOv8-pose and ViTPose pipelines at every temporal threshold, with the margin widening under strict timing. An ablation shows that two intuitively helpful additions---dense foundation-feature change gating and body-part segmentation---both hurt, arguing that a minimal, keypoint-only design is the right one for this task. Finally, standard coaching statistics computed from our automatic predictions track ground truth closely (Pearson r=1.00 for climb time, 0.94 for pace), turning ordinary single-camera video into reliable performance metrics with no instrumentation.
著者のコメント
Accepted at AI2ML Conference 2026 (2nd International Conference on Advancement & Innovation in Artificial Intelligence and Machine Learning)
arXiv ID: 2609.30026 / 要約の誤りについて