身体の違いを反映した両眼視覚データを大量生成
BinoGen: Scaling egocentric binocular data for embodied visual perception and learning
この論文をやさしく読む
ひとことで言うと
見る高さや両眼の配置、移動経路を変えながら、注釈付きの一人称両眼映像を大量に合成する仕組みです。
何に役立つ?
深度推定、物体検出、追跡の学習データづくりに使えます。同じ環境を人型とマウス型の視点から見た対をつくり、身体条件が学習へ与える影響も比較できます。
この研究の面白いところ
2,000万枚超の注釈付き画像を生成し、合成データの追加が実環境の視覚課題を改善したと報告しています。身体別の適応と、両方で使う単一モデルの共同学習を比較しています。
どこまで分かった?
入力要旨は末尾が省略記号で終わっており、提供された範囲からの説明です。具体的な改善幅や追加の結論は読み取れず、人やマウスの知覚そのものを再現した実験ではありません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
身体を伴う視覚知覚は、環境との継続的な関わりを通じて蓄積される、時間的に整合した視覚経験に依存する。しかし、一人称視点の両眼観測を密な注釈とともに大規模に収集することは、依然として高コストで難しい。また、視覚経験は環境だけでなく、視点の高さ、視野、両眼の幾何配置、場面内での移動など、観察者の身体的条件にも左右される。 これらの課題に対し、屋内環境で身体的条件を考慮した一人称両眼視覚経験を大規模に生成する自動化フレームワークBinoGenを提示する。生成的な場面合成、確率的な物体配置、外観のランダム化、確率的な軌跡生成、設定可能な両眼カメラ構成を通じて、環境と観察者の変動を同時にモデル化する。同期した両眼動画に加え、深度マップ、オプティカルフロー、表面法線、意味マップ、物体座標、カメラ姿勢を含む密なマルチモーダル教師情報を生成する。BinoGenを用い、教師あり学習のために2,000万枚を超える注釈付き画像からなるデータセットを構築する。 BinoGenの相補的な2つの用途を実証する。第1に、BinoGenのデータを取り入れると、深度推定、物体検出、動画中の物体追跡を含む実世界の視覚知覚が一貫して改善する。第2に、同じ環境における人間を模した観測とマウスを模した観測の組により、観察者の身体的条件が知覚学習にどう影響するかを条件を制御して調べられる。身体的条件に固有の適応は性能を大きく改善し、共同学習では単一モデルが両方の身体的条件で競争力のある性能を示す。これらの結果は、大規模で制御可能な視覚経験が、身体を伴う知覚を改善できることを示している……。 注:公開されている原文のアブストラクトは、この箇所で省略記号とともに途切れている。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-17(UTC)
- 最新改訂
- 2026-09-17 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-17 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Embodied visual perception relies on temporally coherent visual experience accumulated through continuous engagement with the environment. However, collecting large-scale egocentric binocular observations together with dense annotations remains costly and difficult. Moreover, visual experience is shaped not only by the environment but also by the embodiment of the observer, including viewing height, field of view, binocular geometry, and motion through the scene. To address these challenges, we present BinoGen, an automated framework for generating large-scale, embodiment-aware egocentric binocular visual experiences in indoor environments. BinoGen jointly models environmental and observer variation through generative scene synthesis, probabilistic object instantiation, appearance randomization, stochastic trajectory generation, and configurable binocular camera setups. The framework produces synchronized binocular videos together with dense multimodal supervision, including depth maps, optical flow, surface normals, semantic maps, object coordinates, and camera poses. Using BinoGen, we construct a dataset comprising more than 20 million annotated images for supervised learning. We demonstrate two complementary utilities of BinoGen. First, incorporating BinoGen data consistently improves real-world visual perception, including depth estimation, object detection, and video object tracking. Second, paired human-inspired and mouse-inspired observations from the same environments enable controlled investigation of how observer embodiment affects perceptual learning. Embodiment-specific adaptation substantially improves performance, while joint training enables a single model to perform competitively across both embodiments. Together, these results demonstrate that large-scale, controllable visual experience can improve embodied perception...
arXiv ID: 2609.19881 / 要約の誤りについて