The Machine-Assisted Literature: How AI Took Over Academic Publishing, and Who Is Getting Paid for It
In the spring of 2025, Jacqueline Ewart, a professor of communication at Griffith University in Australia, received a Google Scholar alert for a paper she recognized. She had reviewed it months earlier for the Journal of Radio & Audio Media and recommended rejection: the writing sounded machine-generated, and several references, including a 2017 article on community radio attributed to Ewart herself, could not be found. Ewart had never written such a paper. Yet now, the same manuscript resurfaced in World of Media, a journal published by the journalism faculty at Lomonosov Moscow State University, with only one word changed in the title.
The journal took the paper offline pending investigation. One of its authors told Retraction Watch he had used AI solely to polish the English, blaming poorly indexed Indian repositories for the unverifiable references. He offered no explanation for the phantom article attributed to Ewart.
The episode is small, but it encapsulates the larger story. A machine wrote, or helped write, a paper. A reviewer caught the issue because she was cited. The paper was published anyway, in another journal, with no changes. No one involved in the second publication checked the references. This, in 2026, is the reality of academic publishing in media and communication research, and the numbers are now large enough to describe the phenomenon.
How much of the literature is machine-assisted
Nobody knows exactly, because almost no one discloses. But with each new study, the estimates climb higher, and the methods behind them have become harder to dismiss.
The most cited work comes from Dmitry Kobak’s group. In a 2025 peer-reviewed paper in Science Advances, Kobak and colleagues tracked “excess vocabulary”, words such as delve and underscore, whose frequency spiked after ChatGPT, across 15 million PubMed abstracts. They set a lower bound: at least 13.5 percent of 2024 abstracts had been processed with a language model, with figures exceeding 40 percent in some countries and journals.
In August 2026, the same group posted a preprint applying a more sensitive method to full texts in PubMed Central. Their estimate: 52 percent of papers published in 2024, 77 percent of those in 2025, and nearly nine in ten published in December 2025, show signs of AI-assisted writing. The preprint is not yet peer reviewed; other researchers told Nature the figures may not generalize beyond the corpus, and Kobak himself said he first assumed the calculation was wrong.
A separate 2026 study in PNAS by Kyle Siler of the University of Toronto analyzed 7.3 million full-text articles published by Elsevier, Frontiers, MDPI and PLOS between 2020 and 2025, again using LLM-associated vocabulary, and estimated that 57 percent of 2025 articles showed evidence of LLM influence, up from 12 percent in 2023. The number has been contested: in a letter to PNAS, Chad Topaz and Utsav Bahl argued that a shift in vocabulary cannot be converted into a prevalence rate for LLM use, and Siler, in reply, accepted that the 57 percent is a measure of how far the language has moved rather than a count of articles, while maintaining that the shift itself is real and large. Andrew Gray’s analysis of articles in the Dimensions database found LLM tools likely involved in more than one in ten papers published in 2024, with marker rates for authors in China, South Korea, and Taiwan roughly four times those for the UK or Australia.
These studies measure LLM-associated language: shifts in vocabulary that a population of texts would not show without machine involvement. They can say how much of the literature is AI-influenced or AI-assisted; they cannot say which paper was written by a machine, or how much of any paper was. “Machine-written” in the strict sense, a text generated wholesale, is a subset of unknown size. Surveys of self-reported use and audits of fabricated references measure different things again.
The pattern repeats in the social sciences. A study of 30,000 abstracts from 25 top economics journals found that the share containing AI-associated terms climbed from 2.8 percent in 2023 to 6.7 percent in 2024. Organization Science, a leading management journal, conducted its own audit: submissions have risen 42 percent since ChatGPT launched, AI-heavy manuscripts are harder to read and more likely to be rejected, and over 30 percent of its peer reviews now show detectable AI use, which editors describe as uninformative.
Two things are worth separating. Most of this is not fraud. Wiley’s 2025 ExplanAItions survey found 71 percent of researchers use AI for writing assistance, and Kobak’s team suspects the true figure is even higher. Editing, translating, and drafting with a language model is now routine. The problem is what travels with it when no one checks: fabricated references, invented findings, and a flood of submissions that the peer review system was never built to absorb.
The references that do not exist
Fabricated citations are the most visible symptom of the problem, at once the easiest to catch and the hardest to explain away.
A Columbia University team led by Maxim Topaz audited 2.5 million biomedical papers for nonexistent references; according to Fortune‘s account of the study, published in The Lancet, the rate had risen dramatically since 2023, reaching roughly one paper in 277 in the first seven weeks of 2026. At NeurIPS 2025, one of the most selective conferences in computer science, a University of Chester study documented 100 hallucinated citations in 53 accepted papers, all of which had passed three to five expert reviewers. The Journal of Science Communication devoted an editorial in March 2026 to what its editors call ghost references, noting that peer review was never designed to forensically audit reference lists.
As a result, academics have begun to spot their own names on work they never authored. Pauline Couper, a geography professor at York St John University, recounted on Bluesky reviewing a grant application that cited a nonexistent paper attributed to her. Two education researchers, Charles Hodges and Stephanie Moore, were asked by a colleague for a copy of their 2023 paper on instructional presence in e-learning; the paper was fabricated. Weeks later Hodges was asked to review a book proposal listing a Springer volume he and Moore had supposedly edited, another invention. Retraction Watch reports that tips about fake references have soared since ChatGPT’s debut, and in November 2025 flagged a paper citing an article by its own co-founder that never existed.
Editorial frustration is now out in the open. Udo Schuklenk, editor of Bioethics, described receiving a double-digit number of nonsensical submissions from authors with disposable email addresses. In July 2026, a philosopher writing on Daily Nous examined one such paper and discovered the author didn’t exist: an invented academic with a plausible affiliation and an expanding publication record. Running a simple script over the references, he flagged six fabrications in under a minute, a step the journal’s editorial staff could have taken themselves.
Retractions, paper mills, and the business that feeds them
The retraction record tells the rest of the story. Retraction Watch reported just under 55,000 entries in its database at the end of 2024; by early 2026, the count had exceeded 63,000. The year 2023 remains the record-holder, with over 10,000 articles retracted. An analysis of 2025 retractions by MDPI, using Retraction Watch data, found that 43 percent were linked to compromised peer review or paper mills, and a further 23 percent to citation problems, including irrelevant or nonexistent references. Mass retractions have become routine: one Sage journal, the Journal of Intelligent & Fuzzy Systems, has withdrawn more than 1,500 papers; Wiley’s International Wound Journal retracted 242 in just the first few months of 2025.
Paper mills sell authorship. AI has made their product nearly free to manufacture. What has not changed is the business model that rewards them.
Commercial academic publishing charges authors to publish, charges libraries to read, and pays reviewers nothing. Volume is revenue. RELX, the owner of Elsevier, told investors in February 2026 that its primary research business continues to grow through volume, with article submissions rising strongly across its portfolio; the same report lists research integrity among the principal risks to the business. Wiley reported submissions up 25 percent and output up 11 percent in its last fiscal year, alongside record profit margins. Taylor & Francis reported more than 20 percent growth in submissions in 2025. By Journalology’s count from the Dimensions database, Springer Nature and Wiley each grew article output by about 15 percent last year, compared to market growth of roughly 7 percent.
These companies are extraordinarily profitable. Elsevier’s parent segment runs at a 38 percent adjusted operating margin; Taylor & Francis at 37 percent; Wolters Kluwer’s health division at 32 percent.
Ten of the most prominent academic publishers
Owner, latest reported year, revenue and profit for the closest publishing-related segment. Profit measures differ by company (adjusted operating profit, EBITDA or surplus) and rows are not strictly comparable. A profile, not a league table.
| # | Publisher | Owner | Year | Revenue | Profit | Notes |
|---|---|---|---|---|---|---|
| 1 | ElsevierScientific, Technical & Medical segment | RELX PLCListed in London, Amsterdam and New York. Widely held; no controlling shareholder. Largest holders BlackRock (c. 8–11% across listings) and Vanguard (c. 5%); institutions hold about three-quarters of the shares. | 2025 | £2,714m$3,582m | £1,035madjusted operating profit; 38.1% margin | Segment excludes print from 2025; includes Scopus and other databases. Source |
| 2 | Springer Nature | Holtzbrinck Publishing Group 50.6%; BC Partners 34.8%Listed in Frankfurt since October 2024; free float c. 15%. German regulators cleared Holtzbrinck (the von Holtzbrinck family, Stuttgart) to take sole control in July 2026. | 2025 | €1,926.4mgroup; Research segment €1,517.2m | €543.6madjusted operating profit; Research segment €486.4m | Article output up more than 12%; over 53% of primary research articles published open access. Source |
| 3 | Wolters Kluwer Health | Wolters Kluwer NVListed on Euronext Amsterdam. 100% free float, c. 90% institutional; BlackRock largest at c. 5.7%. | 2025 | €1,596m | €512madjusted operating profit; 32.1% margin | Most revenue is clinical software (UpToDate); journals and books (Lippincott, Ovid) sit in the Learning, Research & Practice unit, about 43% of the division. Source |
| 4 | WileyResearch segment | John Wiley & Sons Inc.Listed on NYSE. Institutions own most of the equity, but the Wiley family (Deborah, Peter, Bradford II, Jesse and others, via EPH LLC) holds about 8.1m of the c. 9m Class B shares, which elect 70% of the board. | FY to Apr 2026 | $1,130mgroup $1,677m | $375madjusted EBITDA, 33.2% margin; group operating income $277m | Submissions up 25%, output up 11%; acquired Emerald Publishing for $452m. Source; proxy statement |
| 5 | Cambridge University Press & Assessment | University of Cambridge | 2024–25 | over £1,000m | c. £200moperating profit (2023–24: £203m) | Combines academic publishing with exams and English-language assessment; publishing share not separately reported. Source |
| 6 | Oxford University Press | University of Oxford | 2024–25 | £796m | £75msurplus from trading (£83m adjusted) | Includes schools and English-language teaching; academic division not separately reported. Source |
| 7 | Taylor & Francisincl. Routledge | Informa PLCListed in London. Over 90% institutional; BlackRock 9.4%, Vanguard 5.5%; no controlling shareholder. | 2025 | £670.8m | £245.7madjusted operating profit; 36.6% margin | Largest publisher in humanities and social sciences, including most media and communication journals. Source |
| 8 | Sage | Sage-SMM TrustFounder Sara Miller McCune transferred control to the trust in 2021. | not disclosed | est. $400–500mlast public estimate, 2021 | not disclosed | Private company; publishes more than 1,000 journals. Source |
| 9 | MDPI | Shu-Kun LinPrivately held, Basel; founded and owned by the chemist Shu-Kun Lin. | 2025 | not disclosedLast self-reported: CHF 191m (2020). An independent study estimated $682m in list-price APC revenue for 2023, before discounts and waivers. | not disclosed | Published 261,576 articles in 2025 from 669,000 submissions; 8,350 staff. Fully author-pays. Source; APC estimate |
| 10 | Frontiers Media | Kamila and Henry MarkramPrivately held, Lausanne; founded and controlled by the two neuroscientists. | 2025 | not disclosedRevenue tracks article output at roughly $1,500 per paper after waivers, by analysts’ estimate. | not disclosed | Fully author-pays; cut 600 of 2,000 staff in 2024 after output fell. Source |
Society publishers such as the American Chemical Society and IEEE are excluded because their publishing revenue is bundled with membership and database income. Shareholdings are from company filings and market data as of 2025–26 and shift continuously. Sage, MDPI and Frontiers do not publish accounts; figures shown for them are external estimates. Compiled by Media Power Monitor, September 2026.
Three types of owners stand behind these publishers. The first is the index fund. RELX, Wolters Kluwer, and Informa have no controlling shareholder; their registers are dominated by institutional giants like BlackRock, Vanguard, State Street, Legal & General, and sovereign funds tracking major indices. These funds don’t read journals, they own publishers as they own every company in the index, focused on margin and buybacks: RELX returned £1.5 billion to shareholders through buybacks in 2025 and plans £2.25 billion in 2026; Informa returned £620 million; Wolters Kluwer spent €1.1 billion repurchasing its own shares.
The second is the family owner. Springer Nature emerged in 2015 from the merger of the von Holtzbrinck family’s Macmillan science division with Springer, then owned by private equity firm BC Partners. Holtzbrinck initially kept 53 percent, and by November 2025 held 50.6 percent, with BC Partners at 34.8 percent. In July 2026, Germany’s competition authority cleared Holtzbrinck to take sole control. Wiley, founded in 1807, remains steered by the Wiley family, who control most of the Class B shares that elect the majority of the board, though institutions hold most of the financial value. Sage belongs to a trust established by founder Sara Miller McCune to remain independent. MDPI is owned by founder Shu-Kun Lin; Frontiers by neuroscientists Kamila and Henry Markram. None of these four private companies publishes accounts, leaving the profits of the author-pays model, the main channel for paper-mill volume, the least visible part of the industry.
The third type of owner is the university. Oxford and Cambridge own the two largest university presses, and their surpluses return to their parent institutions. Unlike the other owners, these universities’ reputations depend directly on the quality of the research they publish.
By Journalology’s estimate from the Dimensions database, eight of these companies (excluding the two university presses) account for roughly 40 percent of all research and review articles published globally, a figure likely to shift as 2025 data are finalized. Their combined adjusted operating profit from the segments above exceeds €3 billion a year. Ultimately, this is paid by university libraries, research funders, and authors, while the labor that makes it possible, peer review, remains unpaid.
Attempts to change the practice
The system is not without dissent, and the dissent is growing.
The oldest form is the walkout. Retraction Watch keeps a running list of mass editorial resignations that passed 50 entries in 2025 and 57 by August 2026. The pattern was set in 2015, when the editors of Elsevier’s linguistics journal Lingua left to found Glossa, a diamond open-access journal with no fees for authors or readers, and repeated in 2023, when the entire boards of NeuroImage and NeuroImage: Reports resigned over Elsevier’s article fees to launch Imaging Neuroscience. The pace has picked up. In December 2025, all editors-in-chief and associate editors of Springer Nature’s Journal of Philosophical Logic resigned to start Philosophical Logic at the Open Library of Humanities. In May 2026, the editors of Springer Nature’s Natural Language Semantics launched a rival journal after, they said, the publisher pressed them to raise annual output by 25 percent.
A second approach changes what review produces. Since January 2023, eLife has made no accept-or-reject decision: every paper it reviews is published as a Reviewed Preprint with the referees’ reports and an editorial assessment attached, and the authors decide whether to revise. Clarivate responded in 2024 by withdrawing the journal’s impact factor. A third pays for the labor. The Company of Biologists’ Biology Open now pays pre-contracted reviewers £220 per manuscript if the review arrives on time and meets editorial standards; in 2025 the scheme cut the time to a first decision from 37.7 working days to 5.5, with no measurable drop in review quality.
The publishers themselves have built defenses. The STM Integrity Hub, a shared screening system run by the publishers’ trade association, now has 49 member organizations, a paper-mill detection tool, and the first cross-publisher check for manuscripts submitted to several journals at once. These tools are aimed to catch fraud after it is submitted. They do not change the fee, the unpaid review, or the reward for volume that make the fraud worth committing.
The Media and Journalism Research Center launched its own attempt in August 2026. According to the center’s description, the Media and Information Research Commons does not publish articles but Verified Research Records: the data, method, findings, and a full log of how the work was done, including every stage where AI was used. Each element is reviewed and citable on its own; an article is optional. Two named reviewers are paid from a single fee calculated from the hours they spend, and their signed reviews are published with the record. Everything is open, and any researcher can use a record’s data and code to write their own article from it.
What sets it apart from the other attempts is the AI protocol at its center. MJRC uses AI extensively, with a dedicated department that processes data, transcribing, classifying, and cross-checking material for its media research, without apology. Its position is that the technology cannot, and should not, be excluded. “I have seen a great deal of AI in the work that crosses my desk, in submissions, in reviews, in our own research,” sdaid Marius Dragomir, MJRC’s director. “It is not going to stop, and it should not: a centre that refuses to use AI for processing data will simply be out-competed by one that does. What we need is a protocol, so that every author states at which stage they used it, how, and who checked the result. And we need to give people a choice. Some work will be AI-produced and openly so; some authors will want their work recognised as human.” The HumanProof Registry exists for the second group. The Commons exists so the first group stops hiding.””
The HumanProof Registry, incubated by MJRC, certifies texts as human-authored based on drafts, version histories, and tracked changes, assessed by paid human reviewers rather than detection software, which the registry considers unreliable and easy to game.
Whether academic publishing will embrace new models of transparency and accountability is far from settled. The forces that fuelled the rise of paper mills, rewarding bulk output, prioritizing publisher profits over peer review, and counting articles as currency, remain unbroken. In the same period that submission volumes soared, the world’s largest publishers reported record profits, while the true scale of machine assistance remains hidden in the footnotes. Until the incentives change, the system’s most dependable protections may still be the vigilance of individual scholars, watching for their names in the citations, and sounding the alarm when something doesn’t add up.
Photo: Pexels

