ドローン向け視覚言語モデルの実行場所を比較
Cloud, Edge, or Split? Profiling Onboard and Split Vision-Language Model Deployment for Drone AI
この論文をやさしく読む
ひとことで言うと
ドローンで視覚言語モデルを動かす場所を、機体上、クラウド、両者への分割から選ぶための比較。
何に役立つ?
ドローンの計算資源や通信条件に合わせて、推論方式を決める際の参考になる。
この研究の面白いところ
遅延だけでなく通信量とエネルギーも測り、画像解像度と回線条件で有利な方式が変わることを示した。
どこまで分かった?
比較は SmolVLM-256M を用いた条件に基づく。要旨に具体的な遅延や消費電力の数値は示されていない。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
視覚言語モデルは、ドローンのような端末が自然言語の指示を用いて視覚的な観測を解釈し、複雑な環境について推論することを可能にする。しかし実用化には難点がある。機体上での推論は計算能力、メモリ、エネルギーの制約を受け、クラウドでの推論には通信遅延、帯域の負担、接続への依存が生じる。これらの制約に対し、資源の限られるドローンと能力の高い遠隔サーバーとの間で推論処理を分割する方法が考えられる。ただし、軽量な視覚言語モデルについて、機体上だけ、クラウドだけ、分割処理という三方式の性能上の兼ね合いは、これまで体系的に調べられていない。本論文は、軽量モデルの代表として SmolVLM-256M を使い、三つの実行方式を比較する。画像解像度とネットワーク条件を変え、推論遅延、計算資源の使用、通信量、エネルギー消費を定量化した。結果は、常に最適な一方式はなく、望ましい方式はネットワーク条件と入力画像の解像度の組合せによって決まることを示す。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-21(UTC)
- 最新改訂
- 2026-09-21 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-21 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Vision-Language Models (VLMs) enable edge devices like unmanned aerial vehicles (UAVs) to interpret visual observations and reason about complex environments using natural-language instructions. However, their practical deployment remains challenging as onboard inference is constrained by limited computational, memory, and energy resources, whereas cloud-based inference introduces communication latency, bandwidth overhead, and dependence on network connectivity. To address these limitations, split computing offers a promising alternative by partitioning VLM inference between the resource-constrained UAVs and more capable remote servers. However, the performance trade-offs among fully onboard, cloud-based, and split-computing architectures for lightweight VLMs have not yet been systematically profiled. This paper benchmarks these three deployment paradigms using SmolVLM-256M as a representative lightweight VLM. We quantify their inference latency, computational resource utilization, communication overhead, and energy consumption across varying image resolutions and network conditions. Our results show that no deployment strategy is universally optimal; instead, the preferred strategy depends on the interaction between network conditions and input image resolution.
著者のコメント
8 pages, 10 figures, and 3 tables. Accepted as a full paper in iEdge 2026
arXiv ID: 2609.25415 / 要約の誤りについて