Papers for

computational chemistry software developers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

TopU-LBVS provides realistic multi-target drug screening benchmark

TopU-LBVS: A Realistic Multi Target Benchmark for Ligand Based Virtual Screening

Abstract: Ligand-based virtual screening (LBVS) is a practical first-pass tool in early-stage drug discovery, but existing benchmarks can overestimate performance through random negatives, easy decoys, limited target coverage, and non-standardized evaluation protocols. We introduce TopU-LBVS, a multi-target benchmark for LBVS under hard-negative screening conditions. Starting from curated ChEMBL~35 bioactivity data, TopU-LBVS covers 93 protein targets across 7 protein classes and constructs target-specific screening libraries with property-matched, structurally similar decoys at a fixed 1:40 active-to-decoy ratio. Libraries contain roughly 400 to 10,000 compounds and are designed to reduce simple physicochemical and nearest-neighbor fingerprint shortcuts. TopU-LBVS provides three fixed protocols. TopU-LBVS-full evaluates ChEMBL$^\ast \rightarrow$ TopU generalization across all 93 targets. TopU-LBVS-low evaluates low-data TopU $\rightarrow$ TopU learning within the hard-negative distribution. TopU-LBVS-mini provides a compact seven-target protocol with a paired random-decoy control that changes only the test decoys, enabling low-cost development and direct measurement of the gap between random ChEMBL$^\ast$ and TopU decoys. Across ten reference baselines spanning fingerprint methods, molecular GNNs, fingerprint hybrids, and modern molecular models, performance under random-decoy evaluation degrades sharply under hard-negative screening. We release data, fixed splits, evaluation code, and baseline implementations for reproducible comparison of future LBVS and molecular representation learning methods. Code and data are available at https://github.com/topu-benchmark/topu-lbvs and https://huggingface.co/datasets/topu-benchmark/topu-lbvs.

Thu 24 SeptMachine LearningArtificial Intelligence
The gist
Early drug discovery often uses virtual screening to find molecules that might interact with proteins, but many existing tests make it look easier than it really is. The authors created TopU-LBVS, a challenging test set with many protein targets and realistic 'hard-negative' molecules that resemble actual drugs but are inactive. They show that popular methods perform worse on this tougher test, suggesting it better reflects real-world difficulties. The benchmark includes clear protocols, data, and code to help developers fairly compare new virtual screening methods.
Open → 2609.29740v1

Complete neural initialization speeds up materials density calculations

Complete Neural Electronic Initialization Accelerates Materials DFT

Abstract: We present the first complete machine learning method for accelerating plane-wave density functional theory (DFT) in materials under the projector augmented wave (PAW) formalism. We formalize seven criteria that a \textit{Complete Neural Electronic Initializer} must satisfy for practical end-to-end PAW DFT acceleration. Applying these criteria to prior work reveals two missing structure-dependent components, augmentation occupancies and spin initialization, that prevent existing methods from providing complete reference-free initialization. Controlled ablations show that omitting these components can eliminate or reverse the acceleration obtained via models that only predict the smooth valence density. We satisfy these missing requirements by introducing AugNet, the first general equivariant model for PAW augmentation occupancies, and the first general spin density model for materials, which predicts the smooth spin-difference density and spin-difference PAW augmentation occupancies using predicted magnetic moments to constrain the global magnetic state. Combined with existing valence density models, these components satisfy all seven criteria and form a fully reference-free electronic initializer for materials DFT, requiring no electronic quantities from a converged target calculation. Our method reduces end-to-end DFT wall time by up to ~25% on unseen structures while preserving converged energies.

Fri 18 SeptMachine Learning
The gist
Calculating the properties of materials at the atomic level is very slow because it needs to figure out how electrons behave. The authors found that previous AI methods for speeding this up missed important details. They created new models that include these missing parts and can start the calculations from scratch without extra information. This new approach speeds up the calculations by about a quarter while keeping results accurate.
Open → 2609.21759v1