arXiv論文メモ
新着一覧
cs.RO · 査読状況未確認

押す動作と音声から飛行船ロボットが人の意図を推定

HINT-Blimp: Human INTent Inference from Multimodal Cues for Robotic Blimps

Subhadeep Koley, Benjamin Greenberg, Yifei Simon Shao, Juan Aceros, Nadia Figueroa, David Saldaña

この論文をやさしく読む

ひとことで言うと

飛行船ロボットに押す動作と声で指示し、ロボットが目的地と動き方を逐次推定する方法を評価した。

何に役立つ?

手持ちの操作器を使いにくい場面で、人の簡単な動作と発話から移動意図を読み取る設計の参考になる。

この研究の面白いところ

300試行で、最大5回のやり取り以内に86%で目標を特定し、多くは2回で済んだ。

どこまで分かった?

検証は飛行船ロボットでの複数参加者による試行で、他のロボットや複雑な環境での性能は要旨に示されていない。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

人とロボットのやり取りで、ジョイスティックや携帯端末のような従来の操作器は、移動作業に遅れを生じさせ、操作者がロボットではなく端末に明示的に注意を向ける必要がある。本研究は、人が押す動作や音声指示といった少数の複数種類の信号で、意図を直接伝える枠組みを提案する。人の意図は、望む目標と動き方を符号化する、パラメータ付き線形力学系として表す。ロボットは粒子フィルターを用いてそのパラメータを逐次推定する。各粒子は候補となる線形力学系の仮説を表し、新しい情報が入るたびに重みを更新する。この枠組みを、物理的な接触を繰り返すのに向いた、しなやかで衝突に耐えやすい飛行船ロボットで検証する。複数の参加者による300試行では、押す動作と音声指示を組み合わせることで、最大5回のやり取り以内に86%の試行で意図した目標を特定し、多くは2回で特定した。推定した力学系は、人だけが知る障害物を避ける曲線状の軌道も作れる。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-22(UTC)
最新改訂
2026-09-22 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

In human-robot interaction, traditional interfaces such as joysticks and handheld tablets introduce latency into navigation tasks and require the operator's explicit attention on the device, instead of the robot. We propose a new human-robot interaction framework in which a human communicates intent directly through sparse multimodal signals such as physical pushes and spoken commands. Human intent is represented as a parameterized linear dynamical system (LDS) that encodes the desired goal and motion behavior. The robot estimates this intent (parameters) online using a particle filter, where each particle represents a candidate LDS hypothesis and is reweighted online as new information becomes available. We validate this framework on a robotic blimp, whose inherent compliance and collision tolerance make it well-suited for repeated physical interaction. Experiments with multiple participants across 300 trials show that combining pushes and voice commands identifies the intended goal in 86% of trials within at most five interactions, with most trials resolved in two. The inferred dynamical systems can also produce curved trajectories that avoid obstacles known only to the human.

著者のコメント

Under review at ICRA 2027

arXiv ID: 2609.27154 / 要約の誤りについて