arXiv論文メモ
新着一覧
cs.RO / cs.AI · 査読状況未確認

MyBuddyロボットに音声対話と知識検索を統合

LLM-based Conversational AI Knowledge Assistant for MyBuddy Humanoid Robot

Hanxiao Chen

この論文をやさしく読む

ひとことで言うと

小型人型ロボットに、音声を聞く、情報を探す、会話を続ける、音声で答える機能を統合した実装です。

何に役立つ?

ロボットを通じた知識案内や会話支援の構成を検討する際に参考になります。感情面の支援は想定する機能であり、心理的な効果が実証されたとの記載はありません。

この研究の面白いところ

LLMだけで完結させず、音声認識、外部の知識取得、対話管理、音声合成をMyBuddy上の対話に組み合わせています。

どこまで分かった?

要旨には利用者実験、応答の正確さ、遅延、比較評価の数値はありません。ロボットの実装報告と、会話や感情支援の効果の検証は区別する必要があります。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

人間中心の用途に向けて人型ロボットの普及と開発が進んでいるが、知的な会話や自然な対話による知識支援を提供する能力は、従来の規則ベースの対話システム、あらかじめ定めた応答、限られた知識源によって制約されている。大規模言語モデル(LLM)は、自然で適応的、かつ文脈を考慮した人間とロボットの相互作用(HRI)を可能にする強力な基盤として登場した。自然な話し言葉を理解し、複雑な質問について推論し、質の高い会話文脈を維持し、知識を豊富に含む応答を生成できるようにすることで、こうした制約に対応する大きな機会をもたらしている。 本研究では、Raspberry Piで動作する13軸のMyBuddy人型ロボット向けに、LLMを利用した多用途の対話型AI知識アシスタントを新たに提案し、実装する。LLMによる言語理解とAI推論に、リアルタイム音声認識、WikipediaやarXivなどのインターネット上の検索・情報源への拡張可能なアクセスによる知識取得、柔軟な対話管理、自然な音声合成を統合する。これにより、より知的な複数ターンの連続会話と、高度な感情面の支援を伴うHRIを可能にする。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-21(UTC)
最新改訂
2026-09-21 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Humanoid robots are increasingly being popular and developed for human-centered applications, yet their ability to provide intelligent conversations and natural interactive knowledge assistance remains constrained by traditional rule-based dialogue systems, pre-defined responses and limited knowledge repositories. Large language models (LLMs) have emerged as a powerful foundation for enabling natural, adaptive, and context-aware Human-Robot Interaction (HRI), which provides a significant opportunity to address such limitations by enabling robots to understand natural speech language, reason over complicated queries, maintain high-quality conversational context, and generate knowledge-rich responses. In this work, we originally present and implement an LLM-based versatile Conversational AI Knowledge Assistant for the Raspberry-Pi-powered 13-Axis MyBuddy humanoid robot, which integrates LLM-driven language understanding and AI reasoning with real-time speech recognition, knowledge retrieval via extensible access of internet engines (e.g., Wikipedia, arXiv), flexible dialogue management, and natural speech synthesis to enable much more intelligent multi-turn continuous conversations and advanced emotional-support Human-Robot Interaction.

著者のコメント

This work has been accepted as poster presentation for NeurIPS 2026 WiML Workshop

arXiv ID: 2609.24742 / 要約の誤りについて