OmniScientist: An Omni-Modal Omni-Discipline AI Scientist
2026-08-13 • Artificial Intelligence
Artificial IntelligenceComputation and Language
AI summaryⓘ
The authors created OmniScientist, an AI system that can handle all steps of scientific research using many types of raw data, like images and formulas, instead of just simplified summaries. It uses different agents to come up with ideas, run experiments, and write papers, making sure the research is new, valid, and well documented. Tested on real scientific cases from various fields, the system successfully produced complete manuscripts directly from raw evidence. Their results show that being able to perceive detailed data throughout the process helps AI make better scientific discoveries.
foundation modelsmulti-modal AIscientific workflow automationdata perceptionresearch lifecyclestatistical validitynovelty screeningmanuscript preparationheterogeneous dataAI agents
Authors
Bobo Li, Hao Fei, Tianjie Ju, Mong-Li Lee, Wynne Hsu
Abstract
Recent advances in foundation models have enabled AI scientists to automate increasingly complete research workflows, from hypothesis generation and code execution to manuscript preparation. Yet workflow coverage alone does not provide access to the full evidence on which scientific discovery depends. Existing systems typically reason over text, code, labels, or precomputed summaries, leaving scientifically decisive spatial, temporal, cross-channel, and procedural relations unavailable to the agent. We introduce OmniScientist, an end-to-end, omni-modal AI scientist that conducts multidisciplinary research directly from heterogeneous raw evidence. A perception layer and 3 autonomous agents for ideation, experiment, and writeup operate within a deterministic pipeline, allowing observations to shape research questions, experimental decisions, and final claims throughout the research lifecycle. By running idea, rigour, and claim checks in code, the system enforces novelty screening, statistical validity, execution provenance, and numerical traceability. We evaluate OmniScientist on 36 real-data cases spanning 5 discipline families, 4 families of scientific evidence, and modalities including images, signals, audio, video, 3-D structures, trajectories, tables, formulae, and graphs. The system completes the full path from raw data to a compiled manuscript in all 36 cases and achieves a mean overall paper score of 6.3 with the reference reasoning backbone. In paired comparisons against a blind variant that receives only precomputed scalar features, direct perception improves all 7 evaluation dimensions and wins 85% of head-to-head judgments. These results show that lifecycle-wide perception is essential for evidence-grounded scientific discovery and provides a practical path toward broadly capable AI scientists.