Papers for
environmental data analysts
Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.
Generative AI shows promise but falls short automating data extraction
Can Generative AI Automate Data Extraction for Meta-Analysis? A Case Study on Intercropping Research
Abstract: Meta-analysis is the synthesis of information from multiple sources to arrive at an overarching conclusion. There is a large need for meta-analysis in agricultural research to synthesize what is known and analyze overarching patterns. Extracting data from published literature is, however, labor-intensive, time-consuming, and tedious, and is impeded by a lack of standardization in research design, units of measurement, and terminology. These challenges are particularly evident in the domain of crop species mixtures, also called intercropping. With the growing capabilities of LLMs, many recent attempts have focused on building systems and tools to automate data collection, yet rigorous assessment against human-labeled ground truth is often missing. In this research, we evaluate three LLM-based approaches---direct zero-shot prompting, a staged workflow, and a multi-agent system---with six open-weight models to extract data from the intercropping literature. The results are evaluated against the manually curated ground truth and through a downstream statistical analysis. Overall, direct zero-shot prompting is the strongest and most consistent approach, achieving the highest mean similarity-adjusted F1 of 0.577, although none of the approaches is close to fully accurate. In the downstream analysis, most model--approach combinations recover the direction of the relationship between the predictor and outcome variables, but do not estimate its magnitude accurately.
Physics informed machine learning can reduce emissions in structural monitoring
From Grey-Box to Green-Box: When can Physics-Informed Machine Learning Reduce Carbon Footprints in Structural Health Monitoring?
Abstract: Machine learning plays an increasingly vital role in engineering, but the corresponding increase in compute time is not without environmental cost. Physics-informed machine learning or "grey-box" models have been developed to overcome some of the limitations of traditional black-box learners, utilising the physical insight that an engineer would have about the structure they are modelling and have shown promising results in the structural engineering field among many others. This work explores whether an additional advantage could be a reduced environmental impact, considering the relationship between training data quantity and training time, linking this duration to carbon emissions from computing. In a structural health monitoring context, four physics-informed machine learning approaches - spanning Gaussian processes and neural networks - are evaluated: residual modelling, input augmentation, hybrid modelling, and constrained learning. The emissions for training each of the models to reach a given error threshold is compared, and in most examples, shown to be lower for the physics-informed models (with input augmented models being an exception). This reduction in training emissions further compounds the environmental savings achieved by collecting and storing less data. Although promising results, we cannot expect a silver bullet and the case studies demonstrate that a trade-off is needed between the increased complexity that comes from introducing physics into a machine learner, against the gain from reduced training data requirements.
Graph-local method improves uncertainty in earth observation treatments
GeoDose-CP: Graph-Local Conformal Inference for Continuous-Treatment Earth Observation
Abstract: Reliable intervention-oriented uncertainty quantification from Earth observation (EO) remains challenging when continuous treatment shifts, spatial dependence, limited support, and satellite-outcome uncertainty must be addressed simultaneously. Existing causal, conformal, and spatial approaches address parts of this problem, but their direct combination does not generally recover the appropriate interventional reference law because candidate reassignment jointly alters treatment likelihood, standardized residuals, and graph-dependent residual likelihood. This study presents GeoDose-CP, a support-aware conformal framework for localized stochastic potential outcomes under continuous or mixed continuous-atomic treatment. Its central methodological contribution is a graph-local target-orbit law that jointly represents intervention-induced treatment shift, the inverse outcome-scale Jacobian, and spatial residual dependence. The framework further provides exact weighted candidate inversion, a scalable sparse approximation with explicit discrepancy accounting, and refusal under inadequate support. Evaluation used controlled known-truth experiments, MineDoseBench, treatment-density sensitivity analysis, external conformal comparators, and a multi-mine New South Wales (NSW) study. In MineDoseBench, GeoDose-CP achieved mean selective coverage of 0.9692 across 27 configurations and a minimum local q0.05 of 0.8951; exact-sparse auditing produced nine inclusion disagreements over 2,700 targets. In the NSW study, the absence of an auditable longitudinal rehabilitation treatment rendered treatment-dependent inference nonoperational rather than forcing inference through a proxy exposure.
New algorithm improves decision-making models for conservation managers
Windowed A-K-MDP
Abstract: Markov decision processes (MDPs) are used to support decision-making in conservation of biodiversity, but policies, even over small state spaces, can be difficult to interpret for conservation managers. K-MDP methods address this problem by building simpler MDPs with at most K abstract states. We show that the previously proposed A-K-MDP algorithm that relies on selecting a discretisation divisor using binary search can skip better abstract states. To fix this issue, we propose Windowed A-K-MDP, an algorithm that generates every distinct feasible partition induced within a declared divisor window and evaluates candidates until reaching the ideal value loss (J = 0) or exhausting the family of candidates. Across 33 K-MDP instances, Windowed improved 25 and tied 8.