arXiv論文メモ
新着一覧
cs.CV / cs.HC · 査読状況未確認

人を映す動画の分析でAI単独と人による確認を比較

Can Vision-Language Models Analyze Human-Centered Video? Mapping Model Capabilities and Human-AI Collaborative Workflows

Xiyuan Shen, Jiuyang Lyu, Seokhyun Hwang, Huanfen Yao, Shwetak Patel, Zhihan Zhang, Jacob O. Wobbrock

この論文をやさしく読む

ひとことで言うと

人を映す研究用動画の注釈で、AIだけの場合と人がAIの結果を確認する場合を比べた研究です。

何に役立つ?

人間中心の動画研究で、どの作業をモデルに任せ、人が確認するかを設計する参考になる。

この研究の面白いところ

人による検証を加えると人単独より精度が高く、時間と費用も減るという結果を、実際の研究論文から導いた課題で測っている。

どこまで分かった?

HNSの平均値は代表的な15課題での結果であり、各課題でモデル単独が人と同等とは限らない。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

動画には、人の行動、相互作用、その場の状況が豊富に記録され、人間中心の研究で人々を理解するための重要な証拠となる。視覚言語モデルによる動画分析は進歩し、人手のかかる作業を自動化する機会が生まれた。しかし、モデルだけで分析できる場合と、信頼できる分析に人の関与が必要な場合の境界は明らかでない。著者らはまず、人間中心の研究での動画分析の実践を特徴付けた。CHI 2026のフルペーパー1,702件を系統的に調べ、動画に注釈を付けた125件を特定した。反復的な分類から、分析目的、視点、現象、必要な推論、注釈の権限という五つの軸の分類体系を導いた。そこに現れる注釈課題を基に、公開データセットから代表的な15課題のベンチマークを作り、汎用視覚言語モデルの能力と限界を調べた。人とモデルの役割分担として、モデル単独、人単独、モデル出力を人が検証する三つの手順を比較した。課題全体の平均ではモデル単独の注釈は人の精度に近づき、HNSは97.0だった。ここで100は人単独の性能を表す。一方、人による検証は最も高い精度を得てHNSは121.5となり、人単独に比べ人の注釈時間を48.9%、金銭的費用を31.3〜44.5%減らした。結果は、現実の人間中心の動画分析課題と現在のモデル能力を結び付け、人とAIの協働で分析の信頼性と効率を高める方法を明らかにする。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-23(UTC)
最新改訂
2026-09-23 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Video provides a rich record of human behavior, interaction, and situated contexts, offering important evidence for understanding people and conducting human-centered research. As vision-language models (VLMs) become increasingly capable of analyzing video, they offer opportunities to automate this traditionally human-intensive process. Yet a central question remains: when can VLMs analyze human-centered video independently, and when does reliable analysis still require human involvement? To address this question, we first characterize video analysis practices in human-centered research. We systematically analyze all 1,702 CHI 2026 full papers and identify 125 that annotate videos. Through iterative coding, we derive a five-dimensional taxonomy spanning analytic purpose, viewpoint, phenomenon, reasoning requirement, and annotation authority. Grounded in recurring annotation tasks captured by this taxonomy, we construct a benchmark of 15 representative tasks from open datasets to map the capabilities and limitations of a general-purpose VLM. We examine the division of labor between humans and VLMs by comparing three annotation workflows: VLM alone, human alone, and human verification of VLM outputs. Across tasks, VLM-alone annotation approaches human accuracy on average (HNS = 97.0, where 100 denotes human-alone performance), demonstrating substantial potential to automate human-centered video analysis. Human verification achieves the highest accuracy (HNS = 121.5) while reducing human annotation time by 48.9% and monetary cost by 31.3%-44.5% relative to human-alone annotation. Our findings connect real-world human-centered video analysis tasks and current VLM capabilities, and clarify how human-AI collaboration can make VLM-assisted analysis reliable and efficient.

arXiv ID: 2609.27327 / 要約の誤りについて