Papers for

clinical trial designers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

NeuronSifter improves planning of CNS interventions by modeling microenvironment dynamics

NeuronSifter: Intervention Planning in CNS Microenvironments

Abstract: Prioritizing central nervous system (CNS) interventions requires predicting how a dose, route, and schedule act on a partially observed microenvironment, then choosing the measurement that would change the decision. Action-conditioned predictors reduce a regimen to an identity token or a scalar exposure, discarding where and when the target is engaged; handing a point estimate to a separate planner then discards the joint uncertainty that makes a measurement worth running. We therefore treat decision quality as a property of the intervention interface, not of controller placement. NeuronSifter compiles regimens into state-conditional target-occupancy fields with support masks, propagates them through microenvironment dynamics with an occupancy-conditioned diffusion operator, and selects measurements by their expected reduction in intervention loss, assimilating typed outcomes into the same posterior. In a declared synthetic Alzheimer's disease (AD) evaluation over 64 paired scenario blocks, occupancy conditioning lowers trajectory continuous ranked probability score from 0.165 to 0.110 and raises intervention ordering accuracy from 0.760 to 0.880, and every paired benchmark contrast remains separated after Holm correction. Decision-directed acquisition attains terminal risk 0.160 against 0.166 for a matched numerical Bayesian experimental design planner, and reaches the target risk at 0.796 $[0.732,0.873]$ of an earlier design control's cost, while the corresponding ratio against the matched planner, 0.963 $[0.907,1.025]$, is not separated from equality; point-state and dependence-ablated interfaces instead raise risk to 0.220 and 0.199, and a full-posterior external controller ties exactly. Published AD trials supply a separate retrospective endpoint bridge.

Mon 28 SeptMachine Learning
The gist
Treating brain treatments as simple numbers hides important details about where and when drugs work in brain tissue. The authors created NeuronSifter, a method that models how treatments affect specific brain microenvironments over time and uses this to better plan interventions. By predicting how a treatment impacts brain cells and deciding the best measurements to improve decisions, NeuronSifter more accurately ranks treatment options. Tests in a simulated Alzheimer's disease setting show it improves prediction accuracy and lowers treatment risk compared to earlier methods.
Open → 2609.35445v1

Minimax policy improves decision making in bernoulli bandit problems

Multi-Armed Bernoulli Bandits via Minimax Single-Arm Stopping

Abstract: We develop an index policy for finite-horizon Bernoulli multi-armed bandits from minimax solutions to single-arm bandit (SAB) problems. Each SAB problem involves choosing between an unknown Bernoulli arm and a known reward. We show that minimizing worst-case regret of SAB problems over all non-anticipative policies admits an exact semi-infinite linear programming formulation. The resulting stopping policies offer a natural way to compare arms: the higher the known reward against which a policy continues sampling, the more promising the unknown arm. We turn this intuition into indices based on cumulative continuation probabilities, with a monotone adjustment and a reward-shortfall cap. By relating index errors to the regret of single-arm stopping policies, we establish a distribution-free regret bound of $4.45\sqrt{KT}+10.75K$ for $K$ arms and horizon $T$. This bound matches the minimax-optimal regret order established in the literature. The guarantee extends to rewards supported on $[0,1]$ through Bernoulli randomization. We also provide a finite-grid implementation with quantified approximation loss. In numerical experiments, the SAB-based index policy achieves lower worst-case regret than every tested benchmark policy across all evaluated numbers of arms and horizons, while closely matching the grid-based MAB minimax policy in the two-arm setting.

Sat 19 SeptMachine Learning
The gist
Choosing the best option among several uncertain choices is a common problem in fields like online advertising or clinical trials. The authors develop a strategy that focuses on single choices and compares them to known rewards to decide when to stop exploring. Their method guarantees good performance across all possible situations and works well when multiple choices are involved. Tests show this strategy outperforms existing methods, especially when dealing with several options over limited trials.
Open → 2609.22690v1

AI generated data can change causal experiment questions and results

When AI Generates Covariates: Causal Typing and Estimand Drift in Sequential Experiments

Abstract: AI-generated covariates from notes, conversations, images, and wearable streams can change the causal question when their roles are left unspecified. A generated feature may represent a treatment version, pre-action state, history, design variable, mediator, outcome proxy, observation process, or intercurrent event; these roles are not interchangeable. We formulate a causal type discipline for sequential experiments: a versioned representation map, a causal role classifier, a claim-status filter, and an estimand lock. The lock fixes a standardized proximal effect before generated covariates enter the analysis. Under audit correctness and standard identification assumptions, admissible role assignments preserve this estimand. We apply the established conditional-covariance characterization of compression bias to substitution of generated representations for design-relevant states. A standardized decomposition separates compression, conditional-law, and standardization drift. Further results cover mediator adjustment, post-action leakage, marker-intervention conflation, outcome-guided discovery, and state-measurement error. Cluster-level orthogonal estimators distinguish empirical and superpopulation targets under repeated sessions and missing outcomes. Simulations show that refinement helps when it retains design-relevant information, whereas design erasure, leakage, and same-data marker selection can produce bias or undercoverage. The framework places causal semantics and claim status before confirmatory inference with generated representations.

Tue 15 SeptArtificial Intelligence
The gist
When artificial intelligence creates data from things like notes or images, it can change what a study is actually measuring, especially in experiments done step-by-step. The authors explain that these AI-created features can play many different roles in a study, and confusing these roles can lead to mistakes in interpreting results. They propose a clear way to classify AI-generated data and keep track of what question the study answers. Their approach helps avoid errors that come from mixing up treatment effects with other types of data changes. They also tested their ideas with simulations showing their methods can improve analysis when the AI data keeps important information.
Open → 2609.17772v1

Evaluation methods guide healthcare large language model testing

A primer on evaluation methods for large language models in healthcare

Abstract: Large language models (LLMs) have a growing range of applications in medicine, and their evaluation is critical for ensuring they provide benefit and not harm. This evaluation can be more challenging than traditional machine learning for many reasons, including probabilistic and open-ended outputs, and behavior that shifts with prompt design and accumulated context. This review covers four key areas of LLM evaluation: principles of study design, statistical methods, capability evaluation and clinical context evaluation. Capability evaluation considers different benchmarks, including multiple-choice, agentic and multi-turn benchmarks, alongside operational metrics like token usage. Clinical context evaluation addresses establishing accuracy of free text outputs, such as human review and LLM-as-a-judge, and clinical trial approaches. Across sections, we describe underlying concepts and potential pitfalls, while emphasizing the importance of aligning evaluation methods with the research question. Together, this article aims to provide a pragmatic basis for designing and executing rigorous evaluations of healthcare LLMs.

Sun 13 SeptComputation and LanguageArtificial Intelligence
The gist
Large language models are increasingly used in medicine, but making sure they work well is tricky because they give complex and changing answers. The authors review important ways to test these models, including study design, statistical methods, and checking their abilities through quizzes and simulated conversations. They also discuss how to judge the clinical accuracy of the models’ free text answers, such as having humans or the models themselves evaluate outputs and running clinical trials. Their goal is to help people design better tests to make sure these AI tools help rather than harm patients.
Open → 2609.14819v1

Medical studies fall behind fast changing clinical ai models

The widening evaluation gap in medical large language model research 2023 to 2026

Abstract: Large language models are superseded every few quarters; clinical evidence takes years. We asked whether medical research is keeping pace with the systems it evaluates. PubMed returned 11,628 records for January 2023 to June 2026 across fourteen clinical domains, growing 45-fold; 2.5% used a randomised, controlled or prospective design. Evaluation lag, from a study's newest named model release to its own publication, widened from 1.33 to 6.08 quarters. Because discontinued models age mechanically, we benchmarked this against a counterfactual holding model composition fixed: migration to newer systems offset only 56% of the drift (95% CI 50-65). Randomised trials evaluated models a median 4.6 quarters older than other designs (P = 3 x 10^-19), yet among studies naming a model still under development no design differed from any other; 62% of randomised trials evaluated a discontinued family. Rigour and currency are in tension, and that tension reflects model selection rather than research timelines.

Thu 10 SeptComputation and Language
The gist
Medical AI models improve very quickly, but studies testing them take much longer to complete. The authors found that evaluations of these models are getting further and further behind the latest versions. Trials that use strict methods tend to test older models than other study types. This means there’s a trade-off between testing the newest AI and using rigorous study designs. The gap comes mostly from which AI models researchers choose to evaluate, not how long studies take.
Open → 2609.11770v1