arXiv論文メモ
新着一覧
cs.CV / cs.RO · 査読状況未確認

圃場条件での深層学習による綿花実の選択的位置検出

Selective Cotton Boll Localization for Robotic Harvesting: Evaluation of Deep Learning Vision Models Under Field Conditions

Thevathayarajh Thayananthan, Xin Zhang, Isuru Laddusinghe Badu, Jonathan Harjono, Glen C. Rains, Beiwen Li, Leonardo M. Bastos, Nuwan K. Wijewardane, Vitor S. Martins

この論文をやさしく読む

ひとことで言うと

畑で綿の実を見つけて輪郭を分け、選択して収穫するロボット向け画像認識を比較しています。

何に役立つ?

自然光や天候が変わる綿畑で、精度と処理速度を両立する知覚モデルを選ぶための評価です。

この研究の面白いところ

1,008枚の注釈付き画像で多数の検出・分割法を比較し、YOLOv12-m-segはAP@0.5が83.7%、1枚20.4 msでした。UR5eと専用把持部で現場実験も行っています。

どこまで分かった?

画像分割指標やマスク面積のR²は、収穫成功率とは別です。要旨には現場試験の総試行数や、条件ごとの収穫成功率は示されていません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

本研究では、綿花を選択的に収穫するロボットのため、深層学習に基づく認識フレームワークを開発・評価した。データセットは、自然光と天候が異なる条件で3台のカメラを用いて収集した、1,008枚の注釈付き圃場画像から成る。物体検出モデルではYOLOv8からYOLOv13までの系列をデフォルト設定で評価し、セグメンテーション性能ではYOLOv8-seg、YOLOv11-seg、YOLOv12-seg、Segment Anything Model(SAM)、SAMv2.1、FastSAM、Recognize Anything Model(RAM)を伴うGrounded-SAMを評価した。 検出モデルでは、GELAN-sが平均適合率(mAP)と推論速度のバランスで最も良好で、mAP86.1%、適合率81.6%、再現率76.6%、F1スコア79.0%、画像1枚あたり平均推論時間42.3 msを得た。直接セグメンテーションモデルでは、YOLOv12-m-segがAP@0.5とFPSのバランスで最も良く、セグメンテーションAP@0.5は83.7%、推論時間は画像1枚あたり20.4 msだった。検出結果をプロンプトにするセグメンテーションでは、GELAN-sが生成した境界ボックスのプロンプトによりSAMとSAMv2.1の綿花実の位置特定が改善し、SAMv2.1 TinyはFastSAMとRAM付きGrounded-SAMを一貫して上回った。 人手で注釈したセグメンテーションマスクとの面積ベース評価では、YOLOv12-m-segのR²は0.966で、GELAN-s+SAMv2.1 Tinyの0.860を上回った。UR5eロボットアーム、カスタムエンドエフェクタ、ZED2iステレオカメラを用いた圃場実験でも、さまざまな信頼度でのリアルタイム綿花実の検出、セグメンテーション、選択収穫にYOLOv12-m-segが有効であることを検証した。これらの結果は、YOLOv12-m-segが効率的な綿花収穫用認識モデルであり、圃場展開に大きな可能性を持つことを示す。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-17(UTC)
最新改訂
2026-09-17 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

This study developed and evaluated a deep-learning-based perception framework for selective robotic cotton picking. The dataset contained 1,008 annotated field images collected using three cameras under varying natural lighting and weather conditions. Object-detection models from the YOLOv8 through YOLOv13 families were evaluated using their default configurations, while segmentation performance was assessed using YOLOv8-seg, YOLOv11-seg, YOLOv12-seg, the Segment Anything Model (SAM), SAMv2.1, FastSAM, and Grounded-SAM with the Recognize Anything Model (RAM). Among the detection models, GELAN-s achieved the most favorable balance between mean average precision (mAP) and inference speed, obtaining an mAP of 86.1%, precision of 81.6%, recall of 76.6%, and an F1-score of 79.0%, with an average inference time of 42.3 ms per image. Among the direct segmentation models, YOLOv12-m-seg provided the most favorable balance between AP@0.5 and FPS, achieving a segmentation AP@0.5 of 83.7% with an inference time of 20.4 ms per image. In the detection-prompted segmentation approach, bounding-box prompts generated by GELAN-s improved the localization of cotton bolls for SAM and SAMv2.1, while SAMv2.1 Tiny consistently outperformed FastSAM and Grounded-SAM with RAM. In the area-based evaluation against manually annotated segmentation masks, YOLOv12-m-seg achieved an $R^2$ value of 0.966, compared with 0.860 for GELAN-s + SAMv2.1 Tiny. Field experiments conducted using a UR5e robotic manipulator, a custom end-effector, and a ZED2i stereo camera further validated the effectiveness of the YOLOv12-m-seg model for real-time cotton boll detection, segmentation, and selective picking under varying confidence levels. These results demonstrate that YOLOv12-m-seg provides an efficient perception model for robotic cotton harvesting and has strong potential for field deployment.

著者のコメント

27 Pages, 19 Figures, 15 Tables

arXiv ID: 2609.19592 / 要約の誤りについて