会話の感情に合わせて人型ロボットの身ぶりを選ぶ
EmoPose: Vision-Language Model Guided Emotion-Aware Gesture Generation for Humanoid Robots
この論文をやさしく読む
ひとことで言うと
人型ロボットが会話の内容に合わせて身ぶりを選び、実機で安全に実行できる形へ変える仕組み。
何に役立つ?
案内や会話をする人型ロボットの身ぶり計画に役立つ可能性がある。
この研究の面白いところ
言語モデルには動作の種類を選ばせ、関節の目標と実行管理はロボット側のライブラリと決定的な処理が担う。
どこまで分かった?
報告された成績はEmoPose-BenchとUnitree G1での29種類の動作など、要旨に記された評価範囲である。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
社会的なやり取りをする人型ロボットは、発話だけでなく身ぶりでも感情や意図を伝える必要がある。ただし、自由な会話から、その機体で実行できる表現力のある動作を作らなければならない。EmoPoseは視覚・言語モデルを使い、言語、会話履歴、必要に応じた画像から、伝達する種類、ライブラリ内の動作、強さ、発話中の位置を含む順序付きの身ぶり計画を選ぶ。ロボット側が管理する動作ライブラリは14自由度の関節目標を定める。Pose Studioは軌道の自動生成、MuJoCoでのプレビュー、新しい動作とモデル向け案内の同期を支援する。ロボット側の決定的な処理が計画を検証し、軌道の構築、予定、待ち行列、割り込みを管理する。これにより、基盤モデルに関節指令そのものを任せず、制御の窓口を変えずに社会的な場面を増やせる。EmoPose-Benchでは、構造化されたGPT-5.5による計画がEasyで98.25±0.52%、全体で76.50±0.54%となり、同じモデルの直接ラベル選択を上回った。会話文脈の使用と複数動作の順序付けも確認した。MuJoCoの標準的な試験群を完了し、実物のUnitree G1で作成した29種類の動作をすべて実現した。実験室での四つの立ち寄り地点を回る実演では、割り込みを伴う説明、カメラに基づく会話、移動を示した。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-20(UTC)
- 最新改訂
- 2026-09-20 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-20 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Socially competent humanoid robots must communicate affect and intent through gesture as well as speech, yet open-ended interaction must become motion that is both expressive and executable on a specific body. This demands semantic flexibility for contextual social intent while preserving deterministic, embodiment-aware robot control. We present EmoPose, a vision-language model (VLM)-guided framework that bridges this gap through an executable semantic interface. Given language, dialogue history, and optional visual context, the VLM selects an ordered gesture plan containing a communicative class, library variant, intensity, and speech anchor. A scalable robot-owned motion library defines the available expressive vocabulary and the source of 14-DoF joint targets. Pose Studio supports automatic trajectory generation, MuJoCo preview, and automatic synchronization of new library entries with the VLM guide; deterministic robot-side modules validate plans, construct trajectories, schedule gestures, and manage queueing and interruption. This division lets the interaction repertoire grow for new social contexts without changing the control interface or delegating raw joint commands to the foundation model. On the EmoPose-Bench, structured GPT-5.5 planning reaches $98.25\pm0.52\%$ on the Easy tier and $76.50\pm0.54\%$ overall, exceeding same-model direct-label prompting. Further tests validate dialogue-context use and ordered multi-action composition. The system completes the nominal MuJoCo suite and realizes all 29 authored variants on the physical Unitree G1. A four-stop laboratory tour demonstrates expressive narration with interruption, camera-grounded dialogue, and navigation.
arXiv ID: 2609.23414 / 要約の誤りについて