ロボットの長い作業で技能の切り替え時点を学習
StageGuard: Learning Stage Transitions for Long-Horizon Robot Tasks via Agentic Distillation
この論文をやさしく読む
ひとことで言うと
ロボットが長い作業の途中で、現在の技能を終え次へ進む時点を学ぶ手法です。大きな教師モデルの説明を小さな視覚言語モデルへ移します。
何に役立つ?
階層的なロボット制御で、段階の切替を効率よく監視するための方法です。手作業で完了判定を設計する負担を減らす狙いがあります。
この研究の面白いところ
教師の推論と実演軌跡から構造化説明を作り、生徒自身の短い説明を使って学習します。予測評価に加え、閉ループ制御と実機でも検証しています。
どこまで分かった?
段階判定の改善とオンライン監視の効率を報告していますが、要旨には改善率や遅延の数値はありません。あらゆる実作業での安全性を保証する結果ではありません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
階層的計画の枠組みは、長期にわたる課題を実行するため、複数のロボット制御方策の技能を組み合わせる。その際、現在の技能をいつ終了し、次の部分課題へ進むかを判断することが不可欠である。既存の方法は、あらかじめ設計した完了信号の判定器に依存することが多いが、実環境での実行に使える判定器を得るのは難しい。大規模な視覚言語モデル(VLM)は強い推論能力を持つものの、その決定境界は本来、課題の完了基準と一致しているわけではない。また、クラウド上での運用と長い推論は大きな遅延を生み、リアルタイム監視を制約する。 正確で効率的な段階移行判断のため、エージェント型の蒸留枠組みStageGuardを提案する。StageGuardは教師モデルの推論と実演軌跡を組み合わせ、部分課題の完了と方策切り替えについて構造化した説明を生成する。軽量な生徒VLMは、この説明を用いて短い自己説明を生成し、それを教師あり微調整に使う。 二つのベンチマークの軌跡で段階移行の予測を評価し、BEHAVIOR-1Kの階層的ロボット制御に組み込んで閉ループでの課題成功を評価する。さらに実機ロボットでも検証する。結果は、効率的なオンライン監視を支えつつ、段階移行の予測を大幅に改善することを示す。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-17(UTC)
- 最新改訂
- 2026-09-17 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-17 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Hierarchical planning frameworks combine skills from multiple robot control policies for long-horizon task execution, where determining when to terminate the current skill and advance to the next subtask is essential. Existing approaches often rely on pre-designed completion signal checkers that are hard to obtain in real-world execution. Large-scale vision-language models (VLMs) offer strong reasoning capabilities, but their decision boundaries are not inherently aligned with task completion criteria, while cloud deployment and lengthy reasoning introduce substantial latency, limiting real-time monitoring. We propose StageGuard, an agentic distillation framework for accurate and efficient stage-transition decisions. StageGuard combines teacher-model reasoning with demonstration trajectories to generate structured explanations of subtask completion and policy switching. A lightweight student VLM uses these explanations to generate compact self-explanations, which are used for supervised fine-tuning. We evaluate stage-transition prediction on trajectories from two benchmarks and assess closed-loop task success through integration into hierarchical robot control on BEHAVIOR-1K, with further validation on real robots. Results show substantial improvements in stage-transition prediction while supporting efficient online monitoring.
著者のコメント
8 pages, 2 figures
arXiv ID: 2609.20791 / 要約の誤りについて