圃場条件での深層学習による綿花実の選択的位置検出
Selective Cotton Boll Localization for Robotic Harvesting: Evaluation of Deep Learning Vision Models Under Field Conditions
この論文をやさしく読む
ひとことで言うと
畑で綿の実を見つけて輪郭を分け、選択して収穫するロボット向け画像認識を比較しています。
何に役立つ?
自然光や天候が変わる綿畑で、精度と処理速度を両立する知覚モデルを選ぶための評価です。
この研究の面白いところ
1,008枚の注釈付き画像で多数の検出・分割法を比較し、YOLOv12-m-segはAP@0.5が83.7%、1枚20.4 msでした。UR5eと専用把持部で現場実験も行っています。
どこまで分かった?
画像分割指標やマスク面積のR²は、収穫成功率とは別です。要旨には現場試験の総試行数や、条件ごとの収穫成功率は示されていません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
本研究では、綿花を選択的に収穫するロボットのため、深層学習に基づく認識フレームワークを開発・評価した。データセットは、自然光と天候が異なる条件で3台のカメラを用いて収集した、1,008枚の注釈付き圃場画像から成る。物体検出モデルではYOLOv8からYOLOv13までの系列をデフォルト設定で評価し、セグメンテーション性能ではYOLOv8-seg、YOLOv11-seg、YOLOv12-seg、Segment Anything Model(SAM)、SAMv2.1、FastSAM、Recognize Anything Model(RAM)を伴うGrounded-SAMを評価した。 検出モデルでは、GELAN-sが平均適合率(mAP)と推論速度のバランスで最も良好で、mAP86.1%、適合率81.6%、再現率76.6%、F1スコア79.0%、画像1枚あたり平均推論時間42.3 msを得た。直接セグメンテーションモデルでは、YOLOv12-m-segがAP@0.5とFPSのバランスで最も良く、セグメンテーションAP@0.5は83.7%、推論時間は画像1枚あたり20.4 msだった。検出結果をプロンプトにするセグメンテーションでは、GELAN-sが生成した境界ボックスのプロンプトによりSAMとSAMv2.1の綿花実の位置特定が改善し、SAMv2.1 TinyはFastSAMとRAM付きGrounded-SAMを一貫して上回った。 人手で注釈したセグメンテーションマスクとの面積ベース評価では、YOLOv12-m-segのR²は0.966で、GELAN-s+SAMv2.1 Tinyの0.860を上回った。UR5eロボットアーム、カスタムエンドエフェクタ、ZED2iステレオカメラを用いた圃場実験でも、さまざまな信頼度でのリアルタイム綿花実の検出、セグメンテーション、選択収穫にYOLOv12-m-segが有効であることを検証した。これらの結果は、YOLOv12-m-segが効率的な綿花収穫用認識モデルであり、圃場展開に大きな可能性を持つことを示す。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-17(UTC)
- 最新改訂
- 2026-09-17 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-17 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
This study developed and evaluated a deep-learning-based perception framework for selective robotic cotton picking. The dataset contained 1,008 annotated field images collected using three cameras under varying natural lighting and weather conditions. Object-detection models from the YOLOv8 through YOLOv13 families were evaluated using their default configurations, while segmentation performance was assessed using YOLOv8-seg, YOLOv11-seg, YOLOv12-seg, the Segment Anything Model (SAM), SAMv2.1, FastSAM, and Grounded-SAM with the Recognize Anything Model (RAM). Among the detection models, GELAN-s achieved the most favorable balance between mean average precision (mAP) and inference speed, obtaining an mAP of 86.1%, precision of 81.6%, recall of 76.6%, and an F1-score of 79.0%, with an average inference time of 42.3 ms per image. Among the direct segmentation models, YOLOv12-m-seg provided the most favorable balance between AP@0.5 and FPS, achieving a segmentation AP@0.5 of 83.7% with an inference time of 20.4 ms per image. In the detection-prompted segmentation approach, bounding-box prompts generated by GELAN-s improved the localization of cotton bolls for SAM and SAMv2.1, while SAMv2.1 Tiny consistently outperformed FastSAM and Grounded-SAM with RAM. In the area-based evaluation against manually annotated segmentation masks, YOLOv12-m-seg achieved an $R^2$ value of 0.966, compared with 0.860 for GELAN-s + SAMv2.1 Tiny. Field experiments conducted using a UR5e robotic manipulator, a custom end-effector, and a ZED2i stereo camera further validated the effectiveness of the YOLOv12-m-seg model for real-time cotton boll detection, segmentation, and selective picking under varying confidence levels. These results demonstrate that YOLOv12-m-seg provides an efficient perception model for robotic cotton harvesting and has strong potential for field deployment.
著者のコメント
27 Pages, 19 Figures, 15 Tables
arXiv ID: 2609.19592 / 要約の誤りについて