arXiv論文メモ
新着一覧
cs.HC / cs.AI · 査読状況未確認

スポーツ映像の生成と理解をつなぐ行動グラフ

Semantic Action Graph: A Shared Representation for Agent Grounding and Human Interpretation of Sports Highlights

Tica Lin, Deepak Chandran, Gauri Jagatap, Chen Chen, Andrea Fanelli, David Gunawan, and Josh Kimball

この論文をやさしく読む

ひとことで言うと

スポーツ映像の出来事を、人にもエージェントにも読める行動グラフで表します。選手、動作、時点、状態などを関係で結びます。

何に役立つ?

ハイライトの生成根拠を視聴者が検索し、確かめ、好みに合わせるための設計です。生成と閲覧を同じ表現に基づかせます。

この研究の面白いところ

映像のフレームへ戻れる時点と、共通の限定語彙を持つ点が特徴です。試作をサッカーファン12人に使ってもらい、検索や解釈に活用されたと報告します。

どこまで分かった?

12人からのフィードバックによる初期的な証拠です。大規模利用や多様な競技での有効性を確定した評価ではありません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

動画のハイライトを選び、説明を付けるために生成エージェントが使われる機会が増えている。しかし通常、それらは構造化されていない表現やフレーム単位の表現を扱うため、視聴者が出力を検証したり、個人の好みに合うよう方向づけたりすることが難しい。 本研究では、スポーツの試合を、実行者、行動、受け手、時点、状態のノードと、それらを結ぶ役割、時間、結果のエッジで表す、軽量な領域スキーマ「意味的行動グラフ」を提示する。このスキーマには、つながった出来事の系列、共有された閉じた語彙、フレーム単位で指定できる時点という3つの重要な性質がある。そのため、説明付きハイライトを組み立てるエージェントのパイプラインと、同じ構造に対して視聴者が検索・点検を行う視覚的インターフェースという、2種類の利用主体を同時に支えられる。 このスキーマを、4モジュールのハイライト生成パイプラインとグラフインターフェースを組み合わせた設計検討用システムSportSAGEとして具体化し、サッカーファン12人からの評価を報告する。参加者は生成されたハイライトと説明の品質に満足し、試合のハイライトの検索、移動、解釈にグラフインターフェースを利用した。これらの結果は、人が読める小さな単一スキーマが、エージェントによる生成の根拠を与えると同時に、人間による解釈を支援できるという初期的な証拠を与える。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-17(UTC)
最新改訂
2026-09-17 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Generative agents are increasingly used to select and narrate video highlights, but they typically operate over unstructured or frame-level representations. Their output is consequently difficult for a viewer to verify and steer toward individual preferences. We present the semantic action graph, a lightweight domain schema that represents a sports match as performer, action, recipient, moment, and state nodes connected by role, temporal, and outcome edges. The schema demonstrates three key properties: 1) connected event sequences, 2) a shared, closed vocabulary, and 3) frame-addressable moments, making it suitable to serve two consumers at once: an agentic pipeline that composes narrated highlights, and a visual interface through which viewers query and inspect the same structure. We instantiate it in SportSAGE, a design probe pairing a four-module highlight pipeline with a graph interface, and report feedback from 12 soccer fans. Participants were satisfied with the quality of the generated highlights and narratives, and used the graph interface to search, navigate, and interpret the match highlights. These results provide early evidence that one small, human-readable schema can ground agent generation and support human interpretation at the same time.

著者のコメント

5 pages, 3 figures, Accepted for publication at IEEE VIS 2026 Workshop on GenAI, Agents, and the Future of VIS

arXiv ID: 2609.20768 / 要約の誤りについて