Effect of research paper mills enough to start affecting LLM query results

Scientific paper mills and AI-generated content are posing a growing threat to research integrity, while retractions reveal only a fraction of the problem, TalTech researcher Anton Sokolov says.
In early September, a Novaator experiment showed how easily co-authorship of a scientific paper can be purchased on the black market from paper mills. With just a few WhatsApp messages, 20 days and $800, it is possible to become an author of a scientific paper ready for publication. But misconduct involving authorship is only one part of a much broader problem concerning the integrity of scientific literature.
After the experiment was published, Anton Sokolov, a doctoral student at Tallinn University of Technology (TalTech) and a researcher at the Tyche Institute, contacted Novaator. As part of his own research, he has taken a closer look at what happens to bogus science once it makes its way into databases. Among the sources Sokolov uses in his research is the Retraction Watch database, which compiles retractions of scientific papers and other cases involving research integrity from around the world.
"If we look at Retraction Watch's public database, there are about 61,000 retracted papers from 2010 to 2025. Currently, about two papers per thousand are retracted each year. The number of retractions shows how much is being cleaned up, not how much garbage remains," Sokolov said.
As part of his research, Sokolov compiled a dataset containing nearly 15,000 retraction records focusing on computer science literature. Analysis of the data showed that violations directly related to authorship accounted for less than 2 percent of cases. "But the small proportion does not mean that authorship is rarely sold. The notice shows what the publisher was able to prove, not how much is actually being sold," he explained.
The latter is precisely what makes the issue difficult for publishers. Such transactions usually take place in private conversations and there is nothing in the published text that unequivocally indicates that authorship was purchased.
For this reason, Sokolov said, retraction notices much more frequently cite problems with peer review. "A single notice usually lists several reasons at once, with peer-review problems cited in about seven out of 10 cases," he added.
Computer-generated content has also become an increasingly serious problem in academia in recent years. According to Sokolov, for example, the most common reason for retractions in computer science and artificial intelligence is computer-generated text, accounting for about one-third of cases.
In 2023, the number of officially retracted scientific papers exceeded 10,000 for the first time. Sokolov said, however, that the figure reveals little about the true prevalence of problematic papers because the vast majority of retracted papers came from large-scale cleanup programs by a handful of academic publishers. "The number of retractions simultaneously measures the problem itself, its detection and classification and a publisher's ability to see cases through to completion — not the actual prevalence of problematic authorship or the use of artificial intelligence," the researcher said.
Cases involving Estonian scientists
Retraction Watch, a website that monitors research integrity, also lists papers co-authored by Estonian researchers. "As of September, there are 13 retractions linked to Estonian institutions and 11 of them involve co-authors from other countries. About half are older cases, involving issues such as duplicate publication or errors in analysis. The rest are more recent, dating from 2023 onward and involve problems with the review of the papers or their data," Sokolov explained.
According to Retraction Watch data, one retraction linked to Estonia is also flagged as involving a paper mill. "But that does not show whether any of the authors here purchased authorship," Sokolov said.
The researcher believes the impact of paper mills could also extend to artificial intelligence systems. Because large language models are trained on material found online and in scientific publications, flawed or manipulated research can also make its way into them. "The risk is real and it has already been measured. In a study published last year, a language model was asked to assess retracted or questionable papers. In more than 6,500 assessments, it did not mention the retraction even once," the researcher said.
The problem is that artificial intelligence may absorb bogus scientific material even before a paper is officially retracted. The publicly available text-screening tool Problematic Paper Screener has flagged tens of thousands of papers as questionable. According to Sokolov's preliminary pilot calculation, more than 20,000 of these papers had not been retracted as of June.
One possible solution, he said, would be to use digital signatures to verify authorship of scientific publications. "At the moment, an author of a scientific paper is simply a name in the text. In Estonia, we use our ID cards to sign contracts and applications; similarly, each author could digitally sign off on their contribution," Sokolov suggested.
He said digital identification of authors and other parties involved in scientific publishing would help increase transparency and make it more difficult to manipulate authorship. "This has not yet been implemented, but we plan to continue this research," he said.
--
Editor: Marcus Turovski












