arXiv論文メモ
新着一覧
cs.RO · 査読状況未確認

間取り図を手がかりに探索して質問に答えるHFLEX-EQA

Hierarchical Floorplan-Guided Vision-Language Exploration for Embodied Question Answering

Albert Gassol Puigjaner, Kostas Alexis

この論文をやさしく読む

ひとことで言うと

未知の屋内環境で質問に答えるロボットが、推定した間取り図を使って探索する方法。

何に役立つ?

必要な部屋を探して情報を集めるロボットの探索計画を改善する研究に役立つ。

この研究の面白いところ

局所的な観測に加え、階層的な場面グラフと推定間取り図で未観測の部屋の種類を推測する。

どこまで分かった?

ベンチマーク評価と四足歩行ロボットでの導入は示すが、要旨は具体的な改善率や屋内環境の範囲を記していない。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

身体を持つエージェントによる質問応答(EQA)では、初めて入る環境を探索し、必要な情報を集め、場面に関する質問に答える必要がある。最近の手法は視覚言語モデル(VLM)と意味地図や場面グラフを組み合わせて探索を導くが、多くは局所的な観測だけで探索を進め、環境の構造に関する事前知識は十分に使われない。本研究は、オンラインの場面グラフ構築、VLMによる計画、意味情報を使った未探索境界の探索、間取り図の事前情報を組み合わせる階層的なEQAの枠組みHFLEX-EQAを提案する。RGB-D観測から階層的な場面グラフと、自由な語彙で物体を表せる占有地図を段階的に構築する。これによりVLMは、場面グラフ、課題に関係する視覚観測、探索履歴、推定した位相的な間取り図を合わせて推論できる。さらに、間取り図と自由語彙の未探索境界の意味情報を用い、課題に関係するがまだ観測していない種類の部屋へ探索を導く、部屋の発見戦略を導入する。OpenEQAとExploreEQAのベンチマークで評価し、実際の屋内環境で四足歩行ロボットへの導入も示す。結果は、VLMによる階層的な計画と構造的な間取り図の事前情報を組み合わせることが、EQA課題に有益であることを示す。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-22(UTC)
最新改訂
2026-09-22 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Embodied Question Answering (EQA) requires an agent to explore a previously unseen environment, gather relevant information, and answer questions about the scene. Recent approaches leverage Vision-Language Models (VLMs) together with semantic maps or scene graphs to guide exploration. However, exploration is typically driven only by local observations, while structural priors about the environment remain largely unused. We propose HFLEX-EQA, a hierarchical EQA framework that combines online scene graph construction, VLM- based planning, semantic frontier exploration, and floorplan priors. The system incrementally builds a hierarchical scene graph and an open-vocabulary occupancy map from RGB-D observations, enabling a VLM to jointly reason over the scene graph, task-relevant visual observations, exploration history, and an estimated topological floorplan. Furthermore, we introduce a room-discovery strategy that leverages the floorplan and open-vocabulary frontier semantics to guide exploration toward semantically relevant yet currently unobserved room types. We evaluate HFLEX-EQA on the OpenEQA and ExploreEQA benchmarks and demonstrate deployment on a quadruped robot in real indoor environments. Our results demonstrate the benefit of combining VLM-based hierarchical planning with structural floorplan priors for the EQA task.

arXiv ID: 2609.26360 / 要約の誤りについて