AI Slop Risks Future Discovery
With apologies to The Simpsons news anchor Kent Brockman: I, for one, welcome our new AI research overlords. Or maybe not so much. Catherine Aiken, Steph Batalis, and Greg Tananbaum, in “Charting a Course for AI in Science” (Issues, Spring 2026), provide a valuable critique of how artificial intelligence is rapidly reshaping the production of science. They describe how institutional incentives pressure researchers to adopt AI tools without adequate understanding, training, or oversight. At the same time, AI is undoubtedly bringing real innovation in research and development.
As the authors imply, this duality can be seen playing out in scholarly publishing. Commercial publishers have integrated AI into their business models, capitalizing on the bundling of scientific content for model training while deploying AI systems for integrity checks. Yet publishers remain highly vulnerable to AI slop. In one recent case, for example, the journal Frontiers published articles with absurd, AI-generated images and hallucinated citations, demonstrating that the industry is in a rat-race against the technology while looking to profit from it.
I applaud the authors for pointing out the connection between AI and the perverse incentives of the publish-or-perish culture. I was surprised, though, that while they discussed the traditional routes for sharing research through journal publication, they overlooked the growing movement—and associated AI challenges—of sharing research through preprints. By facilitating the rapid, open sharing of discoveries, preprint repositories reduce the dependence on commercial gatekeepers and help move the needle in the publish-perish culture.
However, AI is threatening the preprint ecosystem on two fronts. The servers face severe strain from the demands of the AI ingestion machine, eager to gobble up openly licensed content for model training and context updating. And, similar to the publishers but perhaps on a larger scale, AI slop is flooding the field. The scale of this challenge is reflected in the recent decision by the open-access service arXiv to institute a strict one-year ban for authors of manuscripts containing unchecked AI-generated content, such as hallucinated citations or leftover chatbot instructions.
As the share of AI-generated content grows and eventually overtakes human-generated content in the scholarly record, it poses an increased risk to future discovery.
These challenges in scholarly publishing highlight a broader problem in the scientific enterprise: a possible future where innovation implodes. We do not yet know whether researcher reliance on AI to synthesize literature, draft manuscripts, propose hypotheses, analyze and generate data, and conduct reviews will erode the distinct human capacity for creative thought. We do know, however, that as the share of AI-generated content grows and eventually overtakes human-generated content in the scholarly record, it poses an increased risk to future discovery. Frontier models will continue to ingest the growing body of AI-generated scientific outputs. The resulting feedback loop is likely to degenerate into model collapse, where AI systems degrade in quality and lose the ability to provide novel or accurate insights due to being trained on (recursively input) synthetic information.
We’re unlikely to escape the AI research overlords in the long term. In the short term, the scientific community must anchor its principles and practices in open science. Teaching researchers how to responsibly share data, code openly, and write comprehensive AI-use disclosures should be a prerequisite in the training that the authors advocate for.
Christopher Steven Marcum
Chair
PREreview Advisory Committee
Catherine Aiken, Steph Battles, and Greg Tananbaum’s essay got me thinking about a new “pilot test” being conducted by the Centers for Medicare & Medicaid Services. Intended to evaluate the Wasteful and Inappropriate Service Reduction (WISeR) Model, the test will run for six years in six states, including Washington, where I live. The stated goal of WISeR is to reduce fraud, waste, abuse, and medically unnecessary care by using artificial intelligence and other advanced technologies to screen selected Medicare procedures before treatment is authorized.
AI is increasingly being introduced not only into health care and scientific research, but also into such diverse fields as finance and education. In many cases, AI is not merely analyzing data after the fact. It is becoming an active participant in the systems being studied, as it is in the WISeR test. This development raises an important methodological question: How to conduct valid scientific evaluation when the intervention itself changes the behavior being measured?
Economists, ecologists, sociologists, organizational theorists, and systems scientists have long studied reflexive systems. What may be new is the emergence of reflexive systems in which adaptation occurs in response to AI-driven processes whose internal logic is only partially discernible to participants, researchers, and, in some cases, even the developers of the systems themselves.
Once participants understand that an AI system is influencing decisions, they begin adapting their behavior.
Traditional experimental design assumes that researchers can control an intervention while observing outcomes in a relatively stable environment. AI challenges that assumption. Once participants understand that an AI system is influencing decisions, they begin adapting their behavior. Physicians alter documentation practices. Organizations redesign workflows. Consultants develop optimization strategies. Software vendors modify products. Increasingly, participants employ AI tools of their own to anticipate, respond to, and influence AI-driven processes. As a result, the environment under study becomes a dynamic sociotechnical system characterized by continuous adaptation. The intervention affects participant behavior, participant behavior changes the data being generated, and those data may subsequently influence future versions of the intervention.
This creates several methodological challenges. First, causal inference becomes more difficult. Observed outcomes may reflect adaptation to the AI system rather than the effectiveness of the AI system itself. Second, data reliability may deteriorate. Changes in documentation practices, coding behavior, reporting conventions, and AI-assisted content generation may alter the meaning and consistency of the data over time. Third, adaptive feedback loops may emerge. Participants learn how to respond to the system, while the system increasingly operates on data generated by those responses. We must then ask: Are we evaluating the performance of the AI, or the behavior of a system adapting to the presence of AI?
The answer may have profound implications for the future of scientific research itself. And yet as the authors point out, “Researchers and scientists are largely navigating these changes on their own, with current incentives pushing them to adopt AI through hurried learning of new and insufficiently understood tools, methods, and models.”
This is certainly the case with the WISeR pilot test.
Richard V. Badalamente
Kennewick, Washington
As Catherine Aiken, Steph Batalis, and Greg Tananbaum argue, the rapid adoption of artificial intelligence in science, without safeguards, risks eroding public trust in science. But it is not only that AI-enhanced science makes it harder for others to understand or audit it. AI-enhanced science, plagued by hallucinations and other distortions resulting from the use of AI in scholarship and peer review, will also contaminate future training data for large language models (LLMs). This might put academia on a self-reinforcing path toward a decreasingly reliable evidence base, which would spell trouble not only for scientific advancement but also for trust in science as an institution, already under immense public and political scrutiny.
Many ongoing attempts to guardrail the role of AI at various stages of scholarship address symptoms rather than underlying causes, as publishers and funders begin to require or at least encourage AI disclosures. The National Institutes of Health, for instance, recently decided not to consider grant applications that contain “sections substantially developed by AI.” And consensus guidelines about how to report LLM use are emerging. Additionally, preprint servers such as arXiv penalize LLM-introduced errors with one-year author bans.
More transparency is helpful but, by itself, will not solve a problem that is likely to become more pronounced as researchers are increasingly worried about the effects of AI. On the one hand, a 2022 University of Wisconsin-Madison survey found that three in four scientists who published on AI did not think society is “prepared for the potential effects of AI applications.” At the same time, scientists seem largely unconcerned about the corrosive effects that AI will have on reliable science, with a 2025 Nature survey revealing that two-thirds of scientists think it is appropriate to use AI to create a first draft of a research paper.
Ultimately, the most consequential influences of AI-enhanced scholarship might not be felt within the scientific community, but in society writ large. Specifically, the self-reinforcing corrosive effects of AI-enabled research on our larger knowledge infrastructures pose an existential threat to science itself. More importantly, however, they have the potential to undermine public trust in science as society’s best creator, curator, and arbiter of knowledge.
The self-reinforcing corrosive effects of AI-enabled research on our larger knowledge infrastructures pose an existential threat to science itself.
Public trust in AI cannot be engineered solely through technical fixes. Maintaining trust will require visible institutional accountability, governance legitimacy, and sustained societal alignment. More specifically, trust in AI-enhanced science by public and policy audiences rightfully depends on the reliability of research findings over time and across contexts. For now, AI-enhanced research seems to align well with the career interests of individual scientists, who are getting published and cited more, but not necessarily with the long-term health of the scientific enterprise.
To minimize the risk of losing trust among broad sections of the public, scientists need to stop prioritizing speed and productivity over a reliable, robust scientific process, as Aiken and colleagues argue. Recent Pew Research Center data show that science remains among the most trusted institutions, although trust levels vary by partisanship. If we, as a scientific community, end up spiraling toward less and less reliable AI-enhanced science, however, increased public distrust will not be a problem of science communication or science policy; it will fundamentally undermine science’s standing in society. And those effects will be hard to turn back.
Isabelle Freiling
Assistant Professor, Department of Communication
Faculty Affiliate, Scientific Computing and Imaging Institute & Global Change and Sustainability Center
One-U Responsible AI Fellow
University of Utah
Megan K. Taylor
PhD student, Duke University
Honorary Affiliate Researcher, University of Wisconsin-Madison
Dietram A. Scheufele
Investigator, Morgridge Institute for Research
Madison, Wisconsin