Summary
Scientific research is increasingly created and reviewed using AI, which affects both how quickly studies are produced and how they are checked for quality. The authors explain that changes in how AI helps make research also change how AI reviews it, creating a back-and-forth cycle between researchers and reviewers. They looked at many studies and found that faster AI-driven research increases pressure on evaluation methods, which become automated and easier to repeat. But this also creates chances for people to trick these systems, so institutions create new rules and technical tools to keep the process fair. This ongoing 'arms race' changes how science is done and how future research builds on past work.
Generative AIAgentic AIScholarly publishingResearch productionResearch evaluationAutomationEvaluation manipulationInstitutional policyEcosystem feedbackAdversarial co-evolution
Authors
Chenguang Wang, Ming Li, Adebayo Braimah, Chenrui Fan, Tuo Wang, Weijie Guan, Ruiyi Zhang, Tianyi Zhou, Dawei Zhou
Abstract
Generative and agentic AI are reshaping both the production and evaluation of scientific research. These developments are often studied separately, as questions of how AI can produce research and how AI can review it. We argue that this separation misses an increasingly important feature of scholarly publishing: changes on one side alter the incentives, constraints, and behavior of the other. We synthesize 230 scholarly publications and institutional records using a taxonomy of six connected dynamics: production scaling, evaluation automation, evaluation manipulation, defense mechanisms and policy responses, evasion and side effects, and long-horizon ecosystem feedback. The literature shows an emerging progression in which cheaper and faster research production increases pressure on evaluation, AI-mediated evaluation becomes more scalable and repeatable, participants can exploit evaluator regularities, and institutions respond with technical safeguards and policy controls. These responses can in turn induce evasion, redistribute errors and workload, and shape the scholarly records reused by future research and evaluation systems. Evidence is strongest for production and evaluation at scale, reproducible manipulation, and institutional response, while post-policy adaptation and artifact-level long-horizon feedback remain less directly observed. This systems view shifts attention from isolated AI capabilities toward how scholarly actors and AI systems adapt to one another over time.