arXiv論文メモ
新着一覧
cs.CV / cs.RO · 査読状況未確認

高解像度の深度と低解像度の色画像で行う私的情報に配慮した領域分割

Privacy-Preserving Semantic Segmentation from High-Resolution Depth and Ultra-Low-Resolution RGB

Xuying Huang, Swithinraj Moses Daniel, Sicong Pan, Sebastian Houben, Maren Bennewitz

この論文をやさしく読む

ひとことで言うと

細部が見えにくい低解像度の色画像と高解像度の深度情報を組み合わせ、ロボットが場面内の物体領域を認識する方法。

何に役立つ?

考えられる用途は、撮影時の視覚的な私的情報を抑えつつ、移動ロボットの物体目標ナビゲーションに必要な3Dの意味情報を得ること。

この研究の面白いところ

深度で画像の幾何学的な細部を保ち、色画像は低解像度に制限する非対称な入力設計。2Dの認識結果を3Dへ統合している。

どこまで分かった?

私的情報について示されたのは復元可能性の低下であり、漏えいを完全に防ぐ保証ではない。性能評価は要旨に記されたデータセットと実ロボット実験による。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

移動ロボットが日常環境に広がるにつれ、搭載カメラから生じるプライバシー上のリスクが懸念されている。極めて低い解像度(ULR)のRGB画像なら、撮影時点で視覚的な私的情報の露出を減らせるが、それだけでは意味や空間の理解が大きく制限される。そこで本研究は、高解像度(HR)の深度情報とULRのRGB画像を組み合わせ、密な幾何学的情報を保ちつつ、細かい視覚情報を制限する非対称なセンシング設定を導入する。両者の情報量の大きな差に対処するため、HRの幾何学情報で意味理解に向けたRGBの再構成とRGB-D領域分割を導く、統合された2次元の枠組みを提案する。フレーム単位では信頼できる予測が得られても、この非対称な設定で場面全体を一貫して理解するのは難しい。そこで、2次元の意味的特徴を統合して3次元の領域分割を行う、端から端まで学習する2Dから3Dへの処理系列を開発する。 ScanNetでの実験では、私的情報の保護を目指す手法の中で2Dと3Dの領域分割性能が最良となり、SUN RGB-DとSceneNNへの追加学習なしの転移も最も良かった。私的情報を復元できる程度の分析では、提案するHR深度とULR RGBの入力によって機微情報の復元可能性が下がることを示した。実ロボットによる実験では、得られた3Dの意味情報が、指定された物体を目標とする移動に役立つことを示した。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-23(UTC)
最新改訂
2026-09-23 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

As mobile robots become increasingly integrated into everyday environments, privacy risks arising from onboard cameras have become a growing concern. Ultra-low-resolution (ULR) RGB can mitigate visual privacy exposure at the source, but ULR appearance alone substantially limits semantic and spatial understanding. We therefore introduce a privacy-preserving asymmetric sensing setting that combines high-resolution (HR) depth with ULR RGB, preserving dense geometry while restricting fine-grained visual information. To address the severe information imbalance between HR depth and ULR RGB, we propose a joint 2D framework using HR geometry to guide semantic-oriented RGB reconstruction and RGB-D segmentation. Despite reliable frame-level predictions, consistent scene-level understanding remains challenging under the asymmetric HR depth--ULR RGB setting. We therefore develop an end-to-end 2D-to-3D pipeline that consolidates 2D semantic features for 3D segmentation. Experiments on ScanNet show that our method achieves the best 2D and 3D segmentation performance among privacy-preserving approaches and delivers the strongest zero-shot transfer to SUN RGB-D and SceneNN. Privacy recoverability analysis shows that our proposed HR depth--ULR RGB input reduces the recoverability of sensitive data, and real-robot experiments demonstrate the utility of the resulting 3D semantics for object-goal navigation.

著者のコメント

Xuying Huang and Swithinraj Moses Daniel have equal contribution

arXiv ID: 2609.28360 / 要約の誤りについて