透過型XRで人とシステムの見える範囲のずれを調査
Through Human Eyes and Machine Eyes: Understanding View Mismatch in Video See-Through Extended Reality
この論文をやさしく読む
ひとことで言うと
XRヘッドセットの画像に写る範囲と、装着者が実際に見える範囲の違いを調べた研究です。
何に役立つ?
ヘッドセット画像をAIに渡して状況を判断させる仕組みで、見落としや余計な情報の取り込みを評価する際に役立ちます。
この研究の面白いところ
両者の視界を三種類の領域に分け、Meta Quest 3の予備測定と四つの事例で安全性やプライバシーへの影響を検討しています。
どこまで分かった?
可視境界の測定はMeta Quest 3での予備的なものです。四つの事例は潜在的な失敗を示すもので、すべての機器や用途での発生率を示してはいません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
映像透過型の拡張現実(VST XR)システムでは、ヘッドセットのスクリーンショットや取得フレームを、利用者の一人称視点の視覚情報の代わりとして使うことが多い。しかし、システムが取得する視界と利用者に実際に見える範囲は必ずしも一致しない。スクリーンショットは機械が読める長方形のフレームを記録する一方、利用者に見える有効な領域は、より狭く、長方形でない場合がある。本論文は、このVST XRにおける人間とシステムの視界のずれを調べる。両者に見える領域、システムだけに見える領域、人間だけに見える領域を定義し、システムの取得領域と人間の可視領域の関係を形式化する。次にMeta Quest 3で予備的に可視境界を測定し、長方形のスクリーンショットと人間に見えるおおよその境界の間に明確なずれがあることを示す。このモデルに基づき、視界のずれがスクリーンショットによるXRのセンシングと、その後の視覚言語モデルのタスクにどう影響し得るかを分析する。代表的な四つの事例を通じて、プロンプトインジェクション、プライバシー漏えい、人間には見えない情報による偏り、人間には見える情報の見落としという潜在的なリスクや失敗の形を示す。結果は、このずれが単なる幾何学上の現象にとどまらず、AIを組み込んだVST XRシステムの安全性、プライバシー、信頼性に関わり得ることを示している。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-24(UTC)
- 最新改訂
- 2026-09-24 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-24 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Video see-through extended reality (VST XR) systems commonly use headset screenshots or captured frames as proxies for the user's first-person visual context. However, the system-captured view and the user's effective visible field do not necessarily coincide: a screenshot records a rectangular machine-readable frame, whereas the user's effective visible region can be more constrained and non-rectangular. This paper studies this human-system view mismatch in VST XR. We formalize the relationship between the system-captured region and the human-visible region by defining their co-visible, system-only, and human-only regions. \rev{We then conduct a pilot-level boundary measurement on Meta Quest 3, revealing a clear mismatch between the rectangular screenshot frame and the approximate human-visible boundary. Building on this model, we analyze how view mismatch can affect screenshot-based XR sensing and downstream vision-language model tasks. Through four representative case studies, we illustrate potential risks and failure modes including prompt injection, privacy leakage, human-invisible information bias, and missing human-visible information. Our results show that view mismatch is not only a geometric artifact, but can also introduce security, privacy, and reliability concerns for AI-integrated VST XR systems.
arXiv ID: 2609.29173 / 要約の誤りについて