ライブ配信の商品と紹介場面を同時に特定する評価基盤
Grounded Product Understanding in Livestream Videos
この論文をやさしく読む
ひとことで言うと
販売ライブ動画で、商品が何かと、その商品を説明している複数の時間帯を同時に特定する課題を作ります。
何に役立つ?
商品ごとの動画切抜きや、離れた時刻に散らばる説明の収集に役立つ研究です。3000の配信例と3.1万超の商品を収録します。
この研究の面白いところ
商品検索と時間位置推定を別々に測るだけでなく、その対応まで評価します。提案モデルはPair mAP@.3を既存最良の10.13%から21.53%へ高めます。
どこまで分かった?
評価はファッション商品を含む指定のベンチマークです。改善後も完全な商品理解に達したとは言えず、要旨はこの課題の難しさを強調しています。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
電子商取引のライブ配信は、オンラインの消費者に商品を紹介する重要な経路となっている。配信には複数の商品が含まれ、それぞれの情報が異なる時点に分散している。このため、商品を中心に配信を切り抜くといった後段の商品理解への応用には大きな課題がある。こうした応用では、モデルは情報収集のために商品と関連区間を特定する必要がある。しかし、一般的な商品理解の既存ベンチマークは、通常、商品検索と時間区間の特定を別々に評価しており、商品の同一性と時間的な根拠との重要な対応関係は、ほとんど評価されていない。 この制約を解消するため、大規模ベンチマークGPUBを導入する。GPUBは、品質管理された複数時点の時間注釈を持つ3,000件のライブ配信事例と、31,000点を超えるファッション商品のカタログからなる。3つの評価課題を提供する。主課題のGrounded Product Understanding(GPrU)では、ライブ配信動画と候補商品集合から、対象商品を特定すると同時に、それを裏付ける場面を時間的に特定する。商品検索と商品登場場面の特定は、補完的な2つの部分課題として位置付ける。 既存のマルチモーダルモデルを評価すると、GPrUは依然として非常に難しく、最良のベースラインでもPair mAP@.3は10.13%にとどまった。この性能差を縮めるため、統合的な商品理解モデルUniProを開発した。共通のマルチモーダル符号化から、商品に対応し、時間構造を持つ表現を導出することで、GPrUにおけるPair mAP@.3を21.53%に改善し、Joint R@1@.3で37.23%を達成した。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-17(UTC)
- 最新改訂
- 2026-09-17 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-17 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
E-commerce livestreams have emerged as an important channel for presenting products to online consumers, containing multiple products whose information is scattered in different moments. This poses significant challenges for downstream product understanding applications, such as product-centric livestream clipping, where models need to identify the product and its relevant segments for information gathering. However, existing benchmarks for general product understanding typically evaluate product retrieval and temporal localization in isolation, leaving the critical correspondence between product identity and temporal evidence largely unassessed. To address this limitation, we introduce GPUB, a large-scale benchmark comprising 3,000 livestream instances with quality-controlled multi-moment temporal annotations and a catalog of over 31K fashion products. GPUB supports three evaluation tasks: the main task Grounded Product Understanding (GPrU) requires jointly identifying the target product and localizing its supporting moments from a livestream video and a candidate product set; Product Retrieval and Product Moment Localization serve as two complementary subtasks. Evaluation of existing multimodal models shows that GPrU remains highly challenging, with the best-performing baseline achieving only 10.13% Pair mAP@.3. To narrow the performance gap, we further develop UniPro, a unified product understanding model that derives product-aligned and temporally structured representations from shared multimodal encoding, improving Pair mAP@.3 to 21.53% while achieving 37.23% Joint R@1@.3 on GPrU.
arXiv ID: 2609.20508 / 要約の誤りについて