1 in 3 ArXiv Papers Now Read as AI-Written: Study
24 Jul 2026
A new study scoring 12,750 ArXiv papers found that roughly a third now read as machine-written — a trend surfacing just as Substack rolls out its own AI-detection feature for writers and publishers.
What the study found
Researchers sampled about 25 papers per field per month on ArXiv from January 2023 through July 2026, running them through an AI-text detector. Across the full sampling window, about 33% of papers were flagged as machine-written. That share held around 32% in the most recent complete quarter and peaked near 39% in early 2026.
The flagged rate began climbing within months of ChatGPT's release, according to the study. But the increase wasn't uniform: computer science papers had the highest flagged share at roughly 65%, while mathematics sat at the opposite extreme, at just 0.7%.
The detector itself has an 85% recovery rate for AI-written academic text and a 0.4% false-positive rate when tested against a pre-LLM control sample of 200 papers per field. That means it catches most AI text but still misses some, and it occasionally mislabels genuinely human-written work.
Substack's parallel move
Separately, Substack announced on July 22, 2026 that it is integrating AI-detection software Pangram, giving users human-vs-AI estimates for posts, notes, replies, and comments over 100 characters. Writers will be able to add an optional AI author's note to disclose AI use, and publishers can pre-scan their own drafts or flag suspected misclassifications.
Substack CEO Chris Best framed the move as "a good use of AI," arguing software should handle everything except having an idea worth reading, caring about, and sharing. The report notes this echoes a broader pattern: social media platforms already label AI-generated photos and videos, and some music streaming services label or penalize AI-generated tracks.
The catch: precision isn't perfect
The detector's numbers cut both ways. An 85% recovery rate implies roughly 15% of AI-written text could slip through undetected, potentially understating true AI usage. Meanwhile, a 0.4% false-positive rate, applied at scale across thousands of papers, could still mislabel a meaningful number of human-written work as AI-generated — a risk that carries real reputational and career stakes for researchers or writers wrongly flagged. There's also a longer-term reliability question: as AI writing styles evolve, detection tools may be gamed or evaded.
Notably, the report does not identify who conducted the ArXiv study, its authors, or its peer-review status, nor does it clarify how the detector was validated across different academic fields or how it distinguishes AI-assisted editing from fully AI-generated text.
Why founders should care
- Detection demand looks likely to grow. With roughly a third of ArXiv papers flagged and platforms like Substack building disclosure tools, founders in content, publishing, or academic-adjacent spaces may increasingly need to account for AI-authorship signals in their products.
- Precision gaps are the product opportunity — and the risk. The 85% recall / 0.4% false-positive tradeoff suggests current detection technology is good but not great, which could open room for improved or more specialized detection products, though it also means any tool a founder ships or relies on may carry real error rates.
- Field-specific variation may matter more than a single detection score. The gap between computer science (65% flagged) and mathematics (0.7% flagged) suggests generic, one-size-fits-all detection tools may underperform compared to domain-tuned approaches — a possible wedge for founders building vertical-specific verification products.
- Platform integrations could become a distribution channel. Substack's Pangram partnership hints that content platforms may prefer to integrate third-party detection rather than build in-house, which could create partnership opportunities for detection-focused startups.
What's still unclear
Several gaps remain in the available reporting: there's no data on the detector's false-negative rate, no comment from ArXiv, flagged authors, or peer reviewers, and no confirmation of whether the ArXiv study and Substack's Pangram tool rely on the same underlying detection technology. How publishers and academic institutions are responding to the flagged-paper findings also remains undocumented.