arXiv論文メモ
新着一覧
cs.LG · 査読状況未確認

配信中に誰へいつ応答するかを選ぶ支援モデル

Live Assistant: Learning Whether, When, and Whom to Assist in Real-World Live Social Streams

Shujian Gao, Jiamei Yan, Yuchen Yang, Penghao Zhou, Qinglei Wang, Tiehan Fan, Yuan Wang, Zuxuan Wu, Yu-gang Jiang

この論文をやさしく読む

ひとことで言うと

ライブ配信の音声・映像やコメントを見ながら、発言するか、記録するか、誰に何を伝えるかを10秒ごとに選ぶ支援モデル。

何に役立つ?

考えられる用途は、配信者や視聴者への状況に応じた支援。ただし要旨で示すのは収集した配信データでの評価である。

この研究の面白いところ

発言せず観察する選択と非公開の記憶更新を明示的に分け、必要なときだけ応答する。

どこまで分かった?

要旨には評価指標の数値があるが、配信現場での支援効果や長期的な利用者への影響は示されていない。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

ライブ配信は、映像と音声、視聴者の活動、配信者の行動、プラットフォーム上の信号が一緒に変化する、長時間続く対話的な環境である。支援の必要性も配信の進行から生じる。本研究は、自発的な働きかけと役割に応じた支援を行うLive Assistantという枠組みを導入し、配信中のやり取りを「行動するか、いつ行動するか、誰に向けるか、何を伝えるか」という結び付いた四つの判断として定式化する。単一の自己回帰方策が10秒ごとに、生の音声・映像、同期したコメント、ギフト、視聴者の動き、配信ルームのメタデータを受け取り、OBS、MEM、ANSのいずれかを選ぶ。OBSは発言せず、MEMは非公開の意味情報の更新を記録し、ANSは宛先、タスク、根拠に基づくメッセージを指定する。この課題のため、実際の配信セッションを構造化された因果的教師信号へ再構成する軌跡エンジンを作り、320時間を超える最適化用の軌跡と、人手で確認した275本のクリップ、13,812の判断時点からなる評価基準を得た。方策の学習には、まれな構造化された判断を強めるMarker-Aware Multiturn Supervised Fine-Tuning(MA-MSFT)を使い、続けて、自己生成した軌跡をターン単位と軌跡単位の寄与に基づいて最適化するStreaming Multiturn GSPO(SM-GSPO)を用いる。学習に使わなかった評価データでは、状態の正解率71.14、宛先の正解率72.67、タスクの正解率58.41を達成し、代表的なストリーミング手法と一般的なマルチモーダル手法を一貫して上回った。この定式化、評価基準、学習枠組みにより、配信支援を共有された社会的な流れへの選択的参加として位置付ける。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-23(UTC)
最新改訂
2026-09-23 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Livestreams are long-lasting interactive environments where audiovisual content, viewer activity, host behavior, and platform signals evolve together, creating assistance needs that emerge from the stream itself. We introduce \liveassistant, a framework for mixed-initiative, role-conditioned assistance that formulates livestream interaction as four coupled decisions: \textbf{whether to act, when to act, whom to address, and what to communicate}. At each 10-second interval, one autoregressive policy consumes native audio and video with synchronized comments, gifts, viewer dynamics, and room metadata, then selects \textsc{OBS}, \textsc{MEM}, or \textsc{ANS}. \textsc{OBS} remains silent, \textsc{MEM} records a private semantic update, and \textsc{ANS} specifies a recipient, task, and grounded message. To support this task, we build a trajectory engine that reconstructs real livestream sessions into structured causal supervision, yielding over 320 hours of optimization trajectories and a human-reviewed benchmark of 275 clips and 13,812 decision intervals. We train the policy with Marker-Aware Multiturn Supervised Fine-Tuning (MA-MSFT), which strengthens sparse structured decisions, followed by Streaming Multiturn GSPO (SM-GSPO), which optimizes self-generated trajectories with turn- and trajectory-level credit. On the held-out benchmark, \liveassistant reaches 71.14 state accuracy, 72.67 recipient accuracy, and 58.41 task accuracy, with consistent gains over representative streaming and general multimodal baselines. Together, the formulation, benchmark, and training framework establish livestream assistance as selective participation in a shared social stream.

著者のコメント

under review

arXiv ID: 2609.27303 / 要約の誤りについて