Yolo models show limits in cross-field weed detection accuracy
A Multi-Dataset Benchmark of YOLO-Based Weed Detection in Precision Agriculture
Computer Vision and Pattern Recognition
Summary
Detecting weeds in farming fields helps reduce unnecessary herbicide use. The authors tested different sizes of YOLO deep learning models on seven weed detection datasets from diverse farms and conditions. While the models worked well when tested on the same data they were trained on, their accuracy dropped a lot when applied to new, different farms. Training on combined datasets helped but did not fully solve this problem. This study shows the need for more adaptable weed detection systems in agriculture.
What this means in practice
- •For precision agriculture teams: Improve weed management by using YOLO models tailored for specific crop and field conditions to balance detection accuracy and speed.
- •For agricultural robot developers: Design weed sensing systems that combine multi-source training data for better robustness across various farming environments.
Authors
Hristina Zdraveska, Vlatko Spasev, Ivica Dimitrovski, Ivan Kitanovski, Petre Lameski
Abstract
Weed detection is an important component of precision agriculture, enabling site-specific weed management and reducing unnecessary herbicide use. Although deep learning methods have achieved strong results for crop and weed detection, many studies rely on single-dataset evaluation, making it difficult to assess robustness across different agricultural domains. This paper presents a multi-dataset benchmark of deep object detectors for weed detection in precision agriculture, with a focused evaluation of YOLO26 models. We evaluate nano, small, and medium variants on seven public weed-detection datasets covering different crops, weed species, field conditions, acquisition setups, and annotation protocols. The models are compared in terms of detection accuracy, model complexity, inference latency, FPS, and model size. In addition to in-dataset evaluation, we investigate cross-domain generalization using a unified one-class weed setup and evaluate multi-source training using the combined training subsets from all datasets. The results show that YOLO26 achieves strong in-dataset performance, with YOLO26m obtaining the highest average accuracy and YOLO26s providing the best practical accuracy-efficiency trade-off. However, cross-domain performance decreases substantially, with YOLO26s dropping from an average in-domain mAP$_{50:95}$ of 0.603 to 0.148 in the off-domain setting. Multi-source training improves performance on several datasets, but does not fully eliminate domain shift. Overall, the benchmark highlights the importance of dataset diversity, domain similarity, and target-domain adaptation for robust weed detection in real-world precision agriculture applications.