arXiv論文メモ
新着一覧
cs.CV / cs.DC / cs.LG · 査読状況未確認

現場の小型機器で未知の動物種を認識する仕組み

Scout: Open-World Species Recognition on the Edge

Mohammad Mehdi Rastikerdar, Hui Guan, Deepak Ganesan

この論文をやさしく読む

ひとことで言うと

未知の動物種をクラウドの視覚言語モデルに時々確認させ、現場の小型機器にも認識能力を学習させる研究。

何に役立つ?

通信と電力が限られる野生動物のカメラトラップで、未知の種を扱う方法の検討に役立つ。

この研究の面白いところ

すべての画像をクラウドに送らず、種の一覧も事前に用意せず、クラウドの判定を現場モデルへ蓄積する。

どこまで分かった?

3地域・30設置での評価である。初期分類対象外の種の正確度は全画像クラウド送信より低く、エネルギー削減との兼ね合いがある。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

大規模な視覚言語モデル(VLM)は、固定された分類対象以外も認識できるが、計算負荷が大きく、多くの小型端末では動かせない。クラウドへの処理委託でこの能力を使えるものの、すべての画像を送ると、限られた通信帯域と通信エネルギーを消費する。本研究は、計算・電力・帯域が厳しく制限される現場に、VLMの未知の種類を認識する能力をどう持ち込むかを問う。設置時には知られていない動物種にも遭遇するカメラトラップによる野生動物の監視を題材とする。 提案するScoutは、クラウドのVLMを時折呼び出して、新しい種類を小型の現場モデルに教える、自律的な認識システムである。設置場所と、動物が写っていない現地の画像だけを与えると、VLMが同定した各種を、設置場所の条件に応じて継続的に認識する能力へ変える。あらかじめ用意した種の一覧、人手によるラベル付け、手動の調整は不要である。 NVIDIA Jetson Orin Nanoを使い、3地域のカメラトラップ設置30件で評価した。Scoutの正確度は、事前に種の一覧を与えられたモデルとの差が0.1~2.5%に収まった。初期の分類対象にない種では正確度53.7~59.1%で、すべてをクラウドに送る方法の56.5~65.1%と比較して、運用時のエネルギーを59~71%減らした。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-19(UTC)
最新改訂
2026-09-19 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Large vision-language models (VLMs) enable recognition beyond a fixed class set, but their computational demands prevent them from running on many edge devices. Cloud offload makes this capability accessible, but sending every image consumes scarce bandwidth and communication energy. We ask how to bring the open-world recognition capability of VLMs to the edge while operating within tight compute, energy, and bandwidth budgets. Wildlife monitoring provides a natural setting for exploring this question because camera traps encounter species not known at deployment. We present Scout, an autonomous open-world recognition system that invokes a cloud VLM intermittently to teach new classes to a compact edge model. Given only the deployment location and empty site frames, Scout autonomously turns each species identified by the VLM into persistent, site-conditioned recognition capability in a resource-efficient edge model, without a predefined species list, human labeling, or manual tuning. Across 30 camera-trap deployments in three regions on an NVIDIA Jetson Orin Nano, the accuracy of Scout remains within 0.1-2.5% of a model given a predefined species list. On species outside its initial class set, Scout achieves 53.7-59.1% accuracy, compared with 56.5-65.1% for full cloud offload, while using 59-71% less deployment energy.

著者のコメント

Under Review

arXiv ID: 2609.22897 / 要約の誤りについて