Plato-Bio: verification-first biological novelty screening with temporal rediscovery and structural benchmarks

2026-07-27Artificial Intelligence

Artificial Intelligence
AI summary

The authors created Plato-Bio, a software tool that helps biology research by organizing steps like literature review, data analysis, and writing into a clear, traceable process. They fixed some issues that could have caused incorrect results and tested the system on tasks like ranking old studies about fish oil and Raynaud phenomenon, and comparing predicted protein structures to real ones. Their tests show the tool works reliably and keeps track of where uncertainties remain, but they caution it is not proof of new scientific discoveries without further review. The authors emphasize the need for careful validation before making stronger claims.

Large language modelPlato-BioProvenance recordsCitation checksRaynaud phenomenonTF-IDFAlphaFoldProtein structure predictionRMSDReproducible workflows
Authors
Stefan G. Creadore
Abstract
Large language model research agents can connect literature retrieval, analysis code, and manuscript preparation, but coherent output does not establish scientific validity. We developed Plato-Bio, a biology-routed extension of the open Plato/Denario architecture that couples explicit workflow states with provenance records, citation checks, claim-to-evidence links, scoped file writes, and publication gates. A source audit identified and repaired three defects that could distort evaluation: loss of task domain in the default factory, omission of declared method signals from scoring, and evidence sidecars that lacked the drafted-claim denominator. On the current clean revision, the full Python suite completed with 931 passes, six skips, and no failures or errors; targeted biology, genomics, evidence/citation, and adversarial-safety suites likewise completed without failure. We evaluated two narrow use cases. In a frozen historical rediscovery task, independent pre-1986 literature bridges ranked the later-studied relation between fish oil and Raynaud phenomenon first; TF-IDF ranked it second and corpus frequency third. This single curated task measures retrospective ranking, not prospective discovery. In a separate comparison of AlphaFold models with experimental structures for 15 human proteins, 11 targets had high-confidence-core C-alpha RMSD below 1 Angstrom (median 0.501 Angstrom). Four targets exceeded 2 Angstrom, and confidence masking reduced the SUMO1 discrepancy from 16.61 to 2.58 Angstrom over 74 residues. The workflow emitted 27 traceable discrepancy regions, all retained as unvalidated hypotheses. Plato-Bio therefore provides reproducible software contracts and auditable screening baselines; broader claims of agent efficacy or biological novelty require preregistered evaluation, independent review, and prospective validation.