言葉で指定された集団にロボットが合流する位置を予測
Where Should I Join? Robot Group Joining via Language-Guided Goal Prediction
この論文をやさしく読む
ひとことで言うと
ロボットが言葉で指定された集団を見つけ、邪魔になりにくい合流位置と向きを予測する方法である。
何に役立つ?
考えられる用途は、案内ロボットなどが人の集団に加わる場面である。要旨では実機ロボットでも合流を示した。
この研究の面白いところ
集団の構成員を候補群から特定し、人の並び方を手掛かりに複数の合流姿勢を評価する。
どこまで分かった?
実験に含まれるのは会話、列、観客の場面である。すべての社会的状況で安全に合流できるとは要旨に示されていない。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
社会的な場でのロボットの移動は、通常、目的地が指定されており、社会的な慣習を守りながらそこへ到達する問題として扱われる。一方、ロボットが集団に加わるには、その集団の現在の活動や並び方から、どこに入るべきかを予測する必要がある。これは意味の理解を強く要する課題だが、ロボットの盲導犬や自律移動スクーターなどにとって重要な能力である。 本論文は、観測結果と対象集団を表す自然言語の説明を与え、関係する構成員を識別し、社会的に適切な合流位置と向きを予測する、言語に基づくロボットの集団合流を定式化する。対象集団の特定には、再帰的なスペクトル分割で構造化された候補部分集合を作り、言語を条件とする画像・幾何モデルで順位付けする。特定した集団に対しては、人間の集団形成に関する事前知識を使う目標予測器が、ロボットの実行可能な姿勢に対する多峰性のエネルギー・向きマップを生成する。 会話、列、観客の各場面について、集団の大きさ、群衆密度、視覚的な曖昧さを変えた実験では、集団特定は1秒未満の推論で競争力のある正確さを達成し、合流姿勢の予測ではすべての比較手法を上回った。実機ロボットの実験でも、静的なやり取りと動的に変わるやり取りの双方で集団への合流を示した。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-23(UTC)
- 最新改訂
- 2026-09-23 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-23 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Social navigation typically assumes a specified goal and focuses on reaching it while respecting social conventions, whereas robot group joining requires predicting where to join based on the group's real-time activity and formation. This is a highly semantic task, yet an important capability for applications such as robotic guide dogs and autonomous mobility scooters. We formulate language-grounded robot group joining: given an observation and a natural-language description of a target group, the robot identifies the relevant group members and predicts socially compliant joining poses. For grounding, we generate structured candidate subsets through recursive spectral partitioning and rank them with a language-conditioned image--geometry model. Given the grounded group, a goal predictor leverages human-formation priors to produce a multimodal energy--orientation map over feasible robot poses. Experiments on conversations, queues, and audiences across varying group sizes, crowd densities, and visual ambiguities show that our method achieves competitive grounding accuracy with sub-second inference and outperforms all baselines in joining-pose prediction. Real-robot experiments further demonstrate group joining in both static and dynamically changing interactions.
arXiv ID: 2609.28467 / 要約の誤りについて