AI Detection

Biased Against ESL and Students with Disabilities? We Put GPTZero 4o to the Test

Being falsely accused of using AI can be damaging, especially for those developing their writing skills. Here’s what our latest testing reveals.

Anna Gao
· 2 min read
Send by email

Early research and social media posts have claimed that AI detectors are biased against English language learners. Educators have also wondered if detectors incorrectly flag students with disabilities, as developing writers tend to use simpler vocabulary and follow more predictable patterns.

These concerns are reasonable, and the cost of a false positive can be high. For students working hard to learn writing in English, being wrongly accused of using AI can dampen their enthusiasm and confidence. 

At GPTZero, we’ve put a lot of effort into eliminating bias against developing writers. GPTZero 4o is the latest result of that work, so we put it to the test on writing from these students.

The Data

We evaluated GPTZero 4o on PERSUADE2.0, a collection of essays written by students in Grade 6 through 12 across the United States with demographic information. We evaluated essays from two groups of students:

All essays came from PERSUADE2.0's official held-out test split, which GPTZero 4o was not trained on.

Transparent And Confident Evaluation Results

Across 914 essays written by English language learners and 1,172 essays written by students with IEP or 504 plans, GPTZero 4o labels all of them as human-written.

GPTZero 4o not only got all of the answers correct but also did so with perfect confidence.

Taking a closer look, there are 147 student essays at the intersection of ELL and disability, and GPTZero 4o assigned them perfect human confidence.

What This Means For Educators

Fairness is a baseline that an AI detector should meet before it enters the classroom. It has guided GPTZero from inception, and GPTZero is dedicated to building a de-biased AI detector. GPTZero 4o is a great milestone; however, it is not a finish line. We will continue testing for potential biases on writing from the students who are at highest risk of being misjudged, keep sharing what we find, and keep raising the bar of fairness and accuracy with every new model release.