Introducing GPTZero 4o
Our new model offers even more reliability for mixed and modified AI text, with broader training and easier ways to understand results
We are launching GPTZero 4o, O standing for Omni: a model that leads not only in AI detection, but also in mixed and modified AI text – and is versatile across different writing styles, as well as different model types, including Gemini and DeepSeek. In short, it’s our most accurate model yet, and outperforms competitors across major benchmarks.
Notably, it deliberately tackles the harder cases for AI detection: AI-generated text that has been paraphrased or mixed with human writing.
We’ve added features to make results easier to understand, including AI Patterns and Masked Detection.
The most accurate model
GPTZero 4o is state-of-the-art performance in AI detection.
With this model, we scaled up our training data across academic writing and social media writing, and revamped the model architecture to be robust to paraphrasing.
Below are our AI detector accuracy and robustness across three major benchmarks compared to other detectors.
EpochAI (Jaeho L. 2026)
Epoch AI built a benchmark consisting of scientific papers, blogs and short fictional works using frontier LLMs including Claude Opus 4.8, GPT-5.5, and Gemini 3.1 Pro Preview. Beyond direct AI generation from a generic prompt, it also includes a 5-shot style imitation, where the LLMs are given five real writings of the same author and asked to imitate their writing style.
GPTZero 4 achieves 0% FPR and beats Pangram4 on FNR on this benchmark.
Graphite* (Paredes, Druck, Benson & Smith, Graphite, May 2026)
Graphite, a research-driven growth agency, built large studies of AI-generated content, including 55k published English articles from Common Crawl web archive (Jan 2020–Mar 2026) with at least 100 words, to study how much of the live web content is AI-generated. Separately, to measure AI detectors' accuracy, they built a ground-truth set with 6k AI articles generated by GPT-5, Gemini 3.1 Pro, Claude Opus 4.6, and 15k pre-ChatGPT human articles. The table below is as Graphite reported.
DetectRL (Wu et al., NeurIPS 2024)
DetectRL is a benchmark to stress test detectors on domains including academic writing, reviews, news, and social media. AI texts were generated from 4 older LLMs (Llama-2-70b, Claude-instant, ChatGPT, Google-PaLM) and perturbed via 8 various adversarial attacks, such as word substitution, prompt variation, and writing noise, to mimic how people try to evade AI detection. The original paper reports F1 scores across several axes: across domains, LLMs, and attacks, with dedicated human writing and out-of-domain generalization splits.
Here’s the comparison of GPTZero 4o versus Pangram:
Model capabilities
One of the biggest challenges in AI detection, we’ve previously announced, is that AI-generated text is often rewritten.
GPTZero 4o, with its very design, helps tackle the scale of this issue. Here are some updates with 4o we are especially proud of.
Easier to interpret
Our AI patterns feature, available in ‘Advanced Scan’ mode, identifies, extracts and explains recurring stylistic and rhetorical patterns in your document, helping you understand what makes writing look AI-like and why.
An AI Pattern is a content, language or grammatical structure characteristic of AI chatbot writing – for example, arranging ideas in neat sets of threes, or framing claims using a “Not X, but Y” structure. This release includes ten writing patterns, which you can read more about here.

AI writing pattern: Phantom Expert detected in our AI Patterns view for a GPT 5.6-Sol written document. The ‘Phantom Experts’ pattern assigns a claim to unnamed authorities instead of identifying a source.
Punctuation does not determine the result
One of our major advantages is that GPTZero 4o and our models are designed to be indifferent to punctuation changes. This comes from understanding the people who use GPTZero: we know many write with em dashes, and we don’t want to be biased against that.
In other words, something minor like adding or removing a semicolon or em dash is not going to determine whether a text is flagged as human-written or AI-generated.
Detection focused on the main writing
Our Masked Detection feature keeps the focus on the most meaningful content of your document.
Specifically, we currently mask headers, meaning our detector does not consider headings as a feature, helping prevent changes to those headings from affecting the result. A title in itself doesn’t give a detector much to work with, while the main writing gives stronger content to assess.
We’ve also introduced a new highlighting color: gray. Gray spans show which pieces of text are excluded by our AI detector, which means that you can expect even more consistent results.

GPTZero 4o Performance Across Frontier LLM Families
In recent months, several frontier labs released new LLMs. We'll walk through GPTZero 4o’s performance on GPT-5.6, Claude 5, Gemini 3.6, and Grok 4.5.
We created a benchmark spanning 10 domains, including short stories, blog posts, news articles, social media posts, research papers, and five essay types (expository, descriptive, persuasive, admissions, argumentative). AI texts are generated across these 4 LLM families at different reasoning effort levels (low/medium/high) for each, using very challenging prompts that are hard to detect.
With FPR < 0.03% on human source, GPTZero 4o achieved > 99% AI recall on all all four AI families.
How can you access GPTZero 4o?
As of 20 September 2026, the updated model is available to all users and is the default model offered on GPTZero.
Conclusion
GPTZero 4o is our most accurate model yet, and most importantly, it demonstrates how we have a system we’ll keep improving over time. By expanding our training data and testing against challenging benchmarks, we know we can continue to improve how we address the trickier cases of AI-generated text detection.
We're proud of the results here, and we know they're also a baseline for measuring future progress. Our goal is to make each update more accurate, while giving you the strongest possible explanations of what those results mean.
References
- Wu, J., Zhan, R., Wong, D. F., Yang, S., Yang, X., Yuan, Y., & Chao, L. S. (2024). DetectRL: Benchmarking LLM-Generated Text Detection in Real-World Scenarios. Advances in Neural Information Processing Systems 37 (NeurIPS 2024, Datasets and Benchmarks Track). arXiv:2410.23746.
- Duarte, A. V., Tufts, B., Oke, A., Fang, F., Oliveira, A. L., & Li, L. (2026). Sem-Detect: Semantic Level Detection of AI Generated Peer-Reviews. arXiv:2605.21713.
- Epoch AI. (2026, July 15). AI detectors rarely flag human writing, but sometimes miss AI text imitating real authors. Epoch AI Data Insights. https://epoch.ai/data-insights/ai-detectors-false-negatives
- Paredes, J. L., Druck, G., Benson, B., & Smith, E. (2026, May). AI Now Writes as Many Online Articles as Humans. Graphite. https://graphite.io/five-percent/research/ai-now-writes-as-many-online-articles-as-humans-do
- Russell, J., Karpinska, M., & Iyyer, M. (2025). People who frequently use ChatGPT for writing tasks are accurate and robust detectors of AI-generated text. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 5342–5373. arXiv:2501.15654.
- Glickenhaus, B., Thai, K., Russell, J., Masrour, E., Han, Y., Spero, M., & Emi, B. (2026). Pangram 4 Technical Report. arXiv:2607.27183.