arXiv論文メモ
新着一覧
cs.RO / cs.CV · 査読状況未確認

作業支援の記録からロボットの点検視点を学ぶINSPECT

INSPECT: Learning Robot View Selection from Assistant Use

Di Wen, Kailun Yang, Wenhao Guo, Yitian Shi, Junwei Zheng, Yufan Chen, Ruiping Liu, Jiale Wei, Rania Rayyes and Kunyu Peng

この論文をやさしく読む

ひとことで言うと

人が組み立てを確認するときの視点変化を利用し、ロボットが確認しやすい視点を選ぶ方法です。

何に役立つ?

部品の有無や取り付け状態を検査するため、次にどこから見るべきかを決める用途が考えられます。

この研究の面白いところ

候補画像を見ず現在画像と既知の姿勢だけから選び、ギアボックス画像で全確認可能性の人手評価を34.8%から41.7%へ改善します。

どこまで分かった?

評価には注釈付き映像の再生で状態フィードバックを模擬する仕組みを使います。改善後も完全に確認できる割合は限定的です。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

組立品を点検するロボットは、部品が存在するか、正しく取り付けられているかを判断しなければならない。一人称視点による組立支援の際には、頭の動きや作業物の取り扱いによって、これらの確認に必要な証拠が見えるようになり、音声による状態確認が観察と手順上の結果を結びつける。本研究では、部品に関する質問へ答え、次の作業手順を案内するスマートグラス支援システムの記録から、ロボットの視点選好を学習するINSPECTを導入する。 Presence-Invariant TwinSwap(PI-TwinSwap)は、対象の同一性に対する対になった介入を通じて、物体に関する証拠を較正する。主張ごとに対応づけた教師信号は、必要な証拠と、カメラで再現可能な観察の変化を分離する。物体を中心とした較正によって、相対的な視点選好をロボットの姿勢へ適合させ、節単位のスクリーニングによって予測された証拠を確認する。ロボットは、候補視点の画像を使わず、現在の観察と既知の姿勢だけで視点を選ぶ。 評価では、注釈付きの支援動画の再生を用いて状態のフィードバックを模擬し、方策の学習に対象領域の視点ラベルを使用しない。実物のギアボックス組立品の画像では、INSPECTは、比較対象となった正解情報を利用しない方策の中で最も高い視点効用を達成した。現在の視点を維持する場合と比べ、人が評価した完全な検証可能性は34.8%から41.7%へ上昇した。IMPACTに含まれる市販アングルグラインダーの記録では、知覚ヘッドを固定したまま、転移した相対視点選択器によって正しい判断の割合が50.6%から54.3%へ増加した。ソースコードはhttps://github.com/Kratos-Wen/INSPECTで公開されている。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-17(UTC)
最新改訂
2026-09-17 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Robots inspecting an assembly must determine which parts are present and whether they are correctly installed. During egocentric assembly assistance, head motion and workpiece handling reveal evidence for these checks, while spoken state confirmations link observations to procedural outcomes. We introduce INSPECT, which learns robot view preferences from records of a smart-glasses assistant that answers part queries and provides next-step guidance. Presence-Invariant TwinSwap (PI-TwinSwap) calibrates object evidence through paired identity interventions. Claim-indexed supervision separates evidence requirements from camera-reproducible observation changes. Object-centered calibration adapts relative view preferences to robot poses, while clause-level screening checks predicted evidence. The robot selects views using only its current observation and known poses, without candidate images. Evaluation uses annotated assistant-video replay to simulate state feedback, without target-domain view labels for policy training. On images of physical gearbox assemblies, INSPECT achieves the highest view utility among the compared non-oracle policies and raises human-rated full verifiability from 34.8% to 41.7% compared with keeping the current view. On commercial angle-grinder recordings in IMPACT, the transferred relative-view selector increases the correct decision rate from 50.6% to 54.3% with a frozen perception head. The source code is available at https://github.com/Kratos-Wen/INSPECT.

著者のコメント

9 pages, 3 figures, 5 tables. Code: https://github.com/Kratos-Wen/INSPECT

arXiv ID: 2609.20615 / 要約の誤りについて