arXiv論文メモ
新着一覧
cs.CV · 査読状況未確認

農業画像を画素単位で理解するAgriScope

AgriScope: Pixel-Grounded Multimodal Understanding for Agricultural Images

Abderrahmene Boudiaf, Mohamad Alanssari, Irfan Hussain, Sajid Javed

この論文をやさしく読む

ひとことで言うと

農業画像について文章で答えるだけでなく、病変や害虫などが画像のどこにあるかを画素単位で示すモデルです。

何に役立つ?

病害、作物と雑草、害虫、植物種を画像上の位置と結び付けて調べる支援が考えられます。領域説明、領域分割、複数往復の対話を一つの枠組みで扱います。

この研究の面白いところ

50万枚超の画像と1100万の指示応答サンプルを持つAgriGroundを、自動の説明生成・位置対応・マスク生成などを通して構築します。生物学的な意味表現と密な位置表現を統合しています。

どこまで分かった?

要旨は複数の画像言語課題での有効性を報告しますが、具体的な精度値や現場運用の結果は示していません。データとコードの公開は予定として述べられています。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

農業画像の理解には、複雑な実環境の下で、植物の病気、害虫、作物の構造、植物種を細かく認識することが必要である。マルチモーダル大規模言語モデル(MLLM)は近年進歩しているが、既存モデルの出力は依然として文章に限られ、画素単位で内容を画像に対応付ける能力を欠く。本研究では、農業画像理解のために、画素への対応付けを備えた統一的なマルチモーダルの枠組みAgriScopeを導入する。 AgriScopeは画像全体・領域・画素の各レベルの理解を一つの枠組みで同時に支え、農業画像において、画像内の対象に対応付けたキャプション生成、指示表現に基づく領域分割、複数ターンのマルチモーダル対話などを可能にする。生物学的意味の符号化、密な空間表現、画素デコーディングを通じて、生物分野に特化した意味表現と、密な空間的対応付けを統合する。 大規模な対応付け学習を支えるため、農業分野の画素対応付きマルチモーダル指示チューニングデータセットAgriGroundを導入する。これは50万枚を超える画像と1,100万件の指示追従サンプルを含み、植物病害の分析、作物・雑草の識別、害虫の認識、きめ細かな植物学的理解を網羅する。AgriGroundは、マルチモーダルなキャプション生成、句単位の対応付け、領域分割マスクの生成、課題に即した指示の合成を統合する多段階の自動アノテーション工程により構築され、密に対応付けられた教師情報を提供する。 複数の農業向け視覚言語課題での広範な実験により、画素に対応付けたマルチモーダル理解におけるAgriScopeの有効性を示し、農業分野の視覚言語学習と視覚的対応付けに対する有力なベンチマークを確立する。データセットとコードは https://github.com/boudiafA/AgriScope で公開予定である。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-17(UTC)
最新改訂
2026-09-17 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Agricultural image understanding requires fine-grained recognition of plant diseases, pests, crop structures, and botanical species under complex real-world conditions. Despite recent advances in Multimodal Large Language Models (MLLMs), existing models remain limited to text-only outputs and lack pixel-level visual grounding capabilities. In this work, we introduce AgriScope, a unified pixel-grounded multimodal framework for agricultural image understanding. AgriScope jointly supports image-level, region-level, and pixel-level understanding within a unified framework, enabling tasks such as grounded caption generation, referring expression segmentation, and multi-turn multimodal interaction for agricultural imagery. AgriScope integrates biologically specialized semantic representations with dense spatial grounding through biological-semantic encoding, dense spatial representations, and pixel decoding. To support large-scale grounded learning, we introduce AgriGround, a large-scale pixel-grounded agricultural multimodal instruction-tuning dataset containing over 500K images and 11M instruction-following samples spanning plant disease analysis, crop and weed identification, insect pest recognition, and fine-grained botanical understanding. AgriGround is constructed through a multi-stage automatic annotation pipeline that integrates multimodal caption generation, phrase-level grounding, segmentation mask generation, and task-oriented instruction synthesis to produce densely grounded supervision. Extensive experiments across multiple agricultural vision-language tasks demonstrate the effectiveness of AgriScope in pixel-grounded multimodal understanding, establishing a strong benchmark for agricultural vision-language learning and visual grounding. The dataset and code will be made publicly available at (https://github.com/boudiafA/AgriScope)

arXiv ID: 2609.20325 / 要約の誤りについて