Papers for

precision agriculture teams

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Robotic agents combine multimodal sensing for smarter crop monitoring

Recent Advances in Agentic Agri-Robotic Phenotyping: A Perspective Review from Fragmented Multimodal Sensing to Unified PhenoAgent Intelligence

Abstract: This review examines the evolution of plant phenotyping from conventional manual trait measurement to high-throughput, robotic, and artificial intelligence-driven crop monitoring. Despite significant advances in imaging, autonomous platforms, multimodal sensing, and deep learning, current phenotyping systems remain fragmented across sensing modalities, crop traits, growth stages, environments, and management objectives. We therefore frame phenotyping as an integrated \emph{seed-soil-plant-environment-management} (SSPEM) intelligence problem, where crop performance reflects interactions among seed quality, root-zone conditions, plant development, environmental exposure, and management actions. The review synthesizes conventional, high-throughput, robotic, and AI-driven phenotyping approaches, highlighting their capabilities and persistent limitations in temporal integration, multimodal reasoning, biological interpretation, and actionable decision support. Building on this analysis, we introduce a conceptual PhenoAgent framework that extends phenotyping beyond the estimation of isolated traits to evidence-based crop-state interpretation, uncertainty-aware reasoning, and management-oriented support. The PhenoAgent concept primarily brings together scattered advances in phenotyping to deliver insights ranging from detailed to high-level, such as what is happening in the crop, why it might be occurring, what evidence is missing, and what actions or additional measurements should be considered. We also discuss challenges in dataset scarcity, annotation, benchmarking, model generalization, and explainability. By linking multimodal phenotyping with agentic AI and closed-loop decision support, this review outlines a path to interpretable, scalable, and deployment-oriented crop intelligence.

Mon 28 SeptComputer Vision and Pattern RecognitionEmerging Technologies
The gist
Measuring plant growth and health usually involves many separate methods and tools, making it hard to get a complete picture. The authors review how plant monitoring has moved from manual checks to using robots, sensors, and AI, but note that systems still often work in isolated parts. They suggest treating plant monitoring as a holistic task that includes seeds, soil, plants, environment, and farming actions together. Their new idea, called PhenoAgent, aims to unify scattered sensing methods into a smart system that can better understand crops, explain what’s happening, and suggest useful next steps for farmers.
Open → 2609.34567v1

Yolo models show limits in cross-field weed detection accuracy

A Multi-Dataset Benchmark of YOLO-Based Weed Detection in Precision Agriculture

Abstract: Weed detection is an important component of precision agriculture, enabling site-specific weed management and reducing unnecessary herbicide use. Although deep learning methods have achieved strong results for crop and weed detection, many studies rely on single-dataset evaluation, making it difficult to assess robustness across different agricultural domains. This paper presents a multi-dataset benchmark of deep object detectors for weed detection in precision agriculture, with a focused evaluation of YOLO26 models. We evaluate nano, small, and medium variants on seven public weed-detection datasets covering different crops, weed species, field conditions, acquisition setups, and annotation protocols. The models are compared in terms of detection accuracy, model complexity, inference latency, FPS, and model size. In addition to in-dataset evaluation, we investigate cross-domain generalization using a unified one-class weed setup and evaluate multi-source training using the combined training subsets from all datasets. The results show that YOLO26 achieves strong in-dataset performance, with YOLO26m obtaining the highest average accuracy and YOLO26s providing the best practical accuracy-efficiency trade-off. However, cross-domain performance decreases substantially, with YOLO26s dropping from an average in-domain mAP$_{50:95}$ of 0.603 to 0.148 in the off-domain setting. Multi-source training improves performance on several datasets, but does not fully eliminate domain shift. Overall, the benchmark highlights the importance of dataset diversity, domain similarity, and target-domain adaptation for robust weed detection in real-world precision agriculture applications.

Sun 27 SeptComputer Vision and Pattern Recognition
The gist
Detecting weeds in farming fields helps reduce unnecessary herbicide use. The authors tested different sizes of YOLO deep learning models on seven weed detection datasets from diverse farms and conditions. While the models worked well when tested on the same data they were trained on, their accuracy dropped a lot when applied to new, different farms. Training on combined datasets helped but did not fully solve this problem. This study shows the need for more adaptable weed detection systems in agriculture.
Open → 2609.33991v1

Foundation models improve hyperspectral image unmixing with resolution fixes

Benchmarking Hyperspectral Foundation Models for Hyperspectral Unmixing

Abstract: Several foundation models dedicated to hyperspectral images have recently been made available. These models are trained on large unlabeled datasets and exhibit strong performance on many hyperspectral imaging tasks, such as classification or denoising. Nonetheless, their performance for hyperspectral unmixing -- the task of separating mixed spectra of overlapping materials in a hyperspectral image -- remain understudied. This might partly be due to the fact that most of them rely on vision transformer backbones, including patchification, leading to a feature resolution problem. While hyperspectral unmixing already arises from the low resolution of hyperspectral images, this patchification step potentially makes the problem even more ill-posed. Therefore, in this work, we aim to answer two questions: 1) \emph{how do foundation models perform in hyperspectral unmixing?}; 2) \emph{how to tackle the feature-level loss of resolution?} To answer the first question, we benchmark foundation models for unmixing, showing that they can reach state-of-the-art performance on four hyperspectral unmixing datasets. To answer the second question, we compare several feature upsampling approaches and empirically show that using a simple one can lead to high performance results. The code is available at https://gitlab.telecom-paris.fr/ring/hfm-hsu.git.

Wed 23 SeptComputer Vision and Pattern Recognition
The gist
Hyperspectral images capture lots of details about materials but often mix signals from overlapping materials, making it hard to tell them apart. The authors studied modern AI models built for these images to see how well they separate mixed materials. They found these models perform very well but lose detail due to how they process images in chunks. By testing simple ways to restore the lost detail, they improved the models’ ability to separate materials accurately.
Open → 2609.28283v1

Compact hyperspectral camera captures spectra in a single snapshot

Compact Low-Cost Hyperspectral Imaging via Angular-to-Spectral Diversity Conversion

Abstract: Snapshot hyperspectral imaging avoids sequential scanning, but systems that jointly achieve stable reconstruction, low cost, and compact optics remain limited. We present a snapshot hyperspectral imaging system based on angular-to-spectral diversity conversion. A tapered kaleidoscope creates replicated views with distinct incidence directions, and a directly attached birefringent filter converts them into view-channel-dependent spectral transmittances, yielding complementary measurements that better condition the inverse problem for more stable single-shot spectral reconstruction. The system preserves a simple pixel-wise linear model for fast non-learning-based reconstruction and uses only off-the-shelf components without relay optics or cascaded modules. We select the birefringent filter configuration using a condition-number-based criterion and validate the system on both synthetic and real data.

Sun 20 SeptComputer Vision and Pattern Recognition
The gist
Taking detailed color pictures that show many wavelengths usually needs complex and slow machines. The authors built a small, inexpensive camera that captures all spectral data in one shot without moving parts. They use a tapered kaleidoscope to create multiple views and a special filter to separate colors in each view, making the data clearer and easier to decode quickly. This new setup only needs common parts and a straightforward math approach, making it practical for real-world use.
Open → 2609.23619v1

Geometry aware clustering improves detection of overlapping plants in drone images

Combining Object Detection with Geometry-Aware Clustering to Distinguish Overlapping Plants in UAV Imagery

Abstract: Reliable plant-level information from unmanned aerial vehicle (UAV) imagery is important for automated crop monitoring. However, in dense crop canopies, adjacent plants frequently overlap and are detected as a single object, reducing the reliability of plant-level measurements. This study presents a geometry-aware post-detection framework for resolving overlapping plant instances using standard RGB UAV imagery. The framework combines object detection with geometric clustering of plant components. Leaves or branches detected within each bush-level region are represented using two complementary geometric features: component centroids and radial intersection points (RIPs) derived from detected plant structures. K-means and Gaussian mixture models determine whether a detected region contains a single plant or two overlapping plants. Density filtering suppresses spurious radial intersections, and a post-pipeline ensemble combines spatial and directional geometric information. The framework was evaluated using UAV imagery of eggplant and tomato crops under field conditions. Centroid-based clustering achieved an F1-score of 0.89 for eggplant, while the combined centroid-RIP approach achieved the best tomato performance, with an accuracy of 0.80, precision of 1.00, and F1-score of 0.75 using K-means. Density filtering substantially improved RIP-based clustering for tomato. The proposed approach provides a lightweight, modular engineering solution that can be integrated with existing RGB UAV monitoring pipelines without additional depth sensors, pixel-level segmentation, three-dimensional reconstruction, or retraining of the primary bush detector. The results demonstrate that geometric reasoning applied to existing detector outputs can complement deep-learning-based object detection and improve plant-level interpretation in dense agricultural canopies.

Fri 18 SeptComputer Vision and Pattern Recognition
The gist
It can be hard to tell where one plant ends and another begins in pictures taken from drones because plants grow close together and overlap. The authors created a method that looks at the shapes and positions of plant parts inside each detected plant area to figure out if there are really two plants stuck together. Their method uses simple math techniques like clustering points that represent leaf centers and special intersection points. This helps separate plants better without needing complicated extra tools or new training, making plant monitoring more accurate for farmers.
Open → 2609.21304v1

Multimodal AI model improves tomato leaf disease diagnosis accuracy

A Multi-Modal Generative Model for Tomato Disease Leaves Understanding

Abstract: Artificial intelligence for plant disease analysis has advanced from task-specific classifiers to multi-modal models capable of jointly interpreting visual and textual information. However, practical deployment in precision agriculture remains limited because most existing approaches treat disease understanding as isolated prediction tasks, failing to capture the complementary relationships among symptom recognition, severity assessment, and question-driven diagnostic reasoning. In tomato pathology, accurate interpretation of diseased leaves requires more than label prediction; it demands integrating visual symptoms with semantic context to support a comprehensive and explainable understanding. Here, we present SOLAR, a multimodal generative model that understands tomato disease spanning six question-answering tasks. SOLAR learns to align visual features with task-aware language representations by Fusion Expert module based on mixture-of-expert, enabling it to generate contextually relevant answers across diverse diagnostic tasks. By formulating tomato disease analysis as a generative Visual Question Answering (VQA) task, SOLAR provides a flexible framework that supports multi-task inference within a single model while improving performance and cross-task knowledge sharing. We evaluate SOLAR on $41,677$ images, including $216,209$ Question-Answering (QA) pairs to understand tomato leaf disease under both closed and open-ended QA settings. Experimental results show that SOLAR consistently outperforms state-of-the-art vision-only, vision-language, and task-specific models across all tasks, demonstrating superior accuracy, robustness, and multimodal reasoning. These findings highlight the potential of generative multimodal modeling as an effective direction for understanding of plant disease. The code for this study is available at https://github.com/EnalisUs/SOLAR.

Thu 17 SeptComputer Vision and Pattern RecognitionArtificial Intelligence
The gist
Tomato plants can get diseases that make their leaves look sick, and understanding these problems is important for farmers. The authors created a new computer program called SOLAR that can look at pictures of tomato leaves and answer different questions about the disease, like what it is and how bad it is. SOLAR uses both images and language together to give better and more complete answers than previous tools. This means it can help explain the disease symptoms more clearly and support better decisions for plant care.
Open → 2609.19555v1

Fog computing predicts cold storage temperature for fresh produce

Real-World Deployment and Performance Characterisation of Fog-Based Deep Learning for Cold-Chain Temperature Prediction over LoRaWAN

Abstract: Fresh fruits and vegetables (FFVs) are highly perishable, and cold-chain breaks contribute significantly to global food waste. While Machine Learning (ML) can enable proactive intervention, cloud-based inference faces challenges such as latency and data loss. Fog computing addresses these issues but has been tested only in simulation for FFV cold-chain temperature prediction. To the best of the authors' knowledge, this paper presents its first real-world deployment. A fog-deployed LSTM-GRU model predicted cold-room temperature using LoRaWAN sensor data collected from a South African apple cold-storage facility with induced cold-chain breaks. Running entirely on a Raspberry Pi 4 with no cloud dependency, the system generated conditional SHAP explanations only when a break is predicted. The deployed system predicts cold-room temperature with an MAE of 0.2°C at roughly 0.2 kWh per day (0.7 Wh per prediction). Predictions were delivered in under one second (555 ms), dominated by network and messaging rather than computation, with conditional explanations adding modest cost. SHAP consumes 28% more CPU but is well within the hardware's capacity. The model attributes its predictions primarily to temperature, humidity, and their interaction. Critically, the deployment surfaced what simulation cannot: a sensor-triggered single point of failure, alongside genuine resilience, autonomous recovery from infrastructure faults and continued operation through internet loss. These are the first published deployment benchmarks for fog-based temperature prediction in FFV cold chains, establishing that explainable temperature forecasting is feasible on resource-constrained edge hardware. Future work includes asynchronous sensor fusion, commercial cold chain deployment, alternative model architectures, and causal analysis.

Sat 12 SeptDistributed, Parallel, and Cluster ComputingMachine Learning
The gist
Fresh fruits and vegetables need to be kept cold to stay fresh, but sometimes the cold storage systems fail, causing food waste. The authors built and tested a system that runs on a small computer near the storage (called fog computing) to predict temperature changes quickly and explain why. This system works even if the internet connection is lost and avoids relying on distant cloud servers, making temperature monitoring more reliable. Their real-world tests showed the system is accurate and energy efficient, and it can warn staff if the cold chain is broken.
Open → 2609.14036v1

Topology-aware vectorization improves agricultural parcel mapping accuracy

Topologically Consistent Agricultural Parcel Vectorization with Semantic-Guided Diffusion and Topology-Aware Polygonization

Abstract: Agricultural parcel polygons play a fundamental role in geospatial applications such as precision agriculture, land administration, and crop monitoring. Beyond regular polygon geometry and low vertex redundancy, practical parcel maps should avoid topological conflicts and preserve common boundaries between adjacent fields. Yet this requirement remains largely unresolved: segmentation-based methods mainly produce parcel masks or raster boundary cues and rely on heuristic raster-to-vector conversion, instance- and contour-based methods reconstruct parcels independently, and recent vector-oriented methods improve polygon regularity but do not explicitly recover adjacent parcels from a shared topological structure. To address this gap, we propose a semantic-guided diffusion framework for topologically consistent agricultural parcel vectorization. It couples joint edge--vertex latent diffusion with supervised multi-cue conditioning to generate geometrically regularised parcel-boundary and vertex primitives while suppressing false-positive responses. A topology-aware parcel polygon reconstruction method then converts these primitives into regular polygons by reconstructing parcel faces from a common planar graph, enabling adjacent predicted parcels to reuse shared boundaries and avoid mutual interior intrusion. Extensive experiments on the AI4SmallFarms and iFLYTEK datasets evaluate parcel vectorization in terms of pixel-level coverage, geometric fidelity, object-level correctness, and topological consistency. The results show strong and competitive performance, with zero measured intrusion ratio and the highest shared-edge recall, demonstrating the potential of the proposed framework for accurate, regular, and topologically consistent agricultural parcel vectorization.

Mon 7 SeptComputer Vision and Pattern Recognition
The gist
Drawing accurate maps of farm fields is tricky because fields share boundaries that should line up perfectly without gaps or overlaps. The paper presents a new computer method that better traces these field edges as neat polygons that share borders correctly. The approach uses advanced techniques to avoid errors and make sure adjacent farm plots don’t overlap or leave gaps. Tests show this method works well on real farm data to create more reliable and useful maps for farming and land management.
Open → 2609.07520v1