I tested the same text with Clever AI Detector, GPTZero, and ZeroGPT, but each AI detection tool gave me a very different result. Which detector is more reliable, and why do their AI content scores vary so much?
A colleague asked me whether AI detector accuracy claims still mean anything after someone edits the output, and I realized I didn’t have a numbers-based answer. I ended up following the data rather than the usual marketing percentages.
The test gets harder after the first draft
The useful starting point was GEDE, short for Generative Essay Detection in Education, a public research dataset from Lukas Gehring and Benjamin Paaßen at Bielefeld University. It includes more than 900 human essays and over 12,500 essays that were generated or modified by language models, so it covers different degrees of AI involvement rather than just untouched output.
The methods are described in the GEDE research paper, while the GEDE dataset code is public. That matters to me because a benchmark is much easier to take seriously when the underlying material can at least be inspected and reused.
I then found a later comparison that took 600 GEDE texts and ran them through several detectors. The texts represented direct AI writing, rewritten AI, AI-improved writing, and humanized AI. That setup makes more sense than testing only raw model output, since raw output is basically the easiest case.
I do have one caveat. I couldn’t independently confirm who conducted that comparision or whether an outside organization was involved. I’m treating the published scores as reported results, not as a fully audited study.
The spread tells more than the headline
The leading overall result was 99.3% for Clever AI Detector results. More interestingly, it still caught 98.7% of the humanized AI samples. Copyleaks was the nearest alternative in the reported data, while several other well-known detectors lost a lot of accuracy once the writing had been transformed.
That’s the part I’d focus on. Most competent detectors looked good when the input was obvious, untouched AI. Once rewriting or human editing entered the equation, the distance between tools became much larger. So an overall accuracy number can hide the exact failure mode people probably care about most.
I also tested the interface briefly. You paste in text, run the scan, get an AI score, and see highlighted passages that contributed to it. The source said the free Clever AI Detector was available without charge, which makes reproducing a few checks fairly low effort.
Based only on this benchmark, Clever came out first and Copyleaks looked like the strongest backup. I still wouldn’t treat any detector score as proof by itself, especially in an academic or workplace decision, but the category-level results seem much more informative than a single accuracy claim.
Has anyone here reproduced these results with GEDE or seen a similarly structured benchmark using different detectors?
If the result is going to be used to accuse someone of cheating, none of these scores is reliable enough on its own. A 90% AI score does not mean there is a 90% probability the author used AI. Each detector uses its own model, threshold, training data, and definition of “AI-written,” so wildly different percentages are normal. Short passages, formal writing, repeated sentence patterns, and heavy editing can make the disagreement worse.
The benchmark @redadmin posted is interesting, but “AI text caught” only measures half the problem. I’d want to see how often each detector incorrectly flags fully human work under the same conditions. Clever AI Detector performed well on those modified AI samples, but without a comparable false-positive rate, I wouldn’t call it definitively the most reliable overall.
For practical use, test a reasonably long sample and treat the result as a reason to review the writing, not as proof. Draft history, sources, revision records, and whether the person can explain their argument are much stronger evidence than choosing whichever detector gives the highest number.
If you’re comparing the percentages directly, the comparison is already misleading because each tool’s “AI score” measures something different. A 70 from GPTZero is not equivalent to a 70 from ZeroGPT or Clever AI Detector. The benchmark @redadmin shared makes Clever look stronger on edited AI text, but that still does not show how it handles different subjects, non-native English, or formulaic human writing. I’d judge a detector by its false positives on writing similar to yours, not by whichever interface displays the most confident number.
Expect these tools to disagree, and don’t assume the highest score comes from the best detector. A benchmark’s “caught” rate can change dramatically depending on the cutoff chosen, so 99% detection means little unless the threshold, raw outputs, and exact detector version are available.
Clever AI Detector looks strongest in the posted test, but the benchmark’s unclear provenance makes me cautious about treating it as the winner. Detector models can be updated without much notice, too, meaning the tool tested months ago may not be identical to the tool you use today.
For your own use, make a small control set of known human and known AI writing from the same subject and similar length. Run that set through all three. The most useful detector is the one that stays consistent on material resembling yours, not the one displaying the boldest percentage. Even then, I’d use the result only as a screening signal.
Running the same draft through several detectors can make the situation worse, because people start editing toward the meters instead of improving the writing. A sentence gets changed until GPTZero approves it, then ZeroGPT dislikes the revision, and soon the document is awkward while nobody has learned anything useful about who wrote it.
The scores vary because the tools are not measuring a shared quantity. Each detector breaks the text into chunks differently, weighs language patterns differently, and applies its own cutoff before displaying a percentage. Some scores refer to the estimated amount of flagged text. Others behave more like confidence ratings. That makes comparing “82%” across three sites a bit like comparing grades from three classes with different exams and grading scales.
Input handling can create another large difference. Quotations, assignment instructions, headings, citations, reference lists, standard disclaimers, and repeated template language may be included by one detector and effectively ignored by another. Even paragraph order and sample length can matter. Before comparing tools, I would remove material the author did not compose and test the same clean body text in each. I would then check longer sections separately rather than repeatedly submitting isolated sentences, since short samples tend to produce jumpier results.
Based on the benchmark posted by @redadmin, Clever AI Detector looks much better at catching AI text that has been rewritten or humanized. That is a useful strength, especially if edited AI is what you are trying to find. I still would not translate that result into “Clever is 99.3% reliable.” The reported number appears to describe how much AI material it caught in that selected test. Reliability in actual use depends on whether it can preserve that sensitivity without flagging an unacceptable amount of human work.
There is a practical way to read disagreement between the three: treat it as an inconclusive result. Do not average the percentages, and do not take a two-out-of-three vote. The detectors may rely on overlapping signals, so three outputs are not necessarily three independent opinions. If Clever flags a passage and the other two do not, that could mean Clever catches subtler editing. It could just as easily mean it applies a more aggressive threshold to that kind of prose.
For low-stakes screening, I would choose the detector that performs best on examples matching the documents being checked. For academic enforcement, hiring, publication disputes, or anything else that can harm someone, none should be the decision-maker. Use the highlighted passages to decide what deserves a closer look, then examine drafts, document history, source use, factual understanding, and revision behavior. The detector can point at smoke, but its percentage cannot tell you with certainty who started the fire.
A detector can look accurate on a full document and still be useless if its verdict flips when you remove the bibliography or split the text in half. That kind of instability is easy to miss when everyone compares only the final percentages.
I would compare Clever AI Detector, GPTZero, and ZeroGPT by running a small sensitivity check. Submit the same clean body text, then submit it again without headings and quotations, and once more as two longer sections. You are not looking for identical numbers, but the classification should remain reasonably consistent. If a tool jumps from “mostly human” to “mostly AI” after harmless formatting changes, its confident-looking score does not deserve much weight.
From the benchmark @redadmin posted, Clever appears strongest at detecting edited or humanized AI, while GPTZero and especially ZeroGPT missed much more of that material. That gives Clever an advantage for that specific task, but it could also mean Clever uses a more aggressive threshold. Without equally clear results for ordinary human writing, “more sensitive” cannot automatically be translated into “more reliable.”
So I would not compare 80% versus 40% as though they were readings from the same instrument. Compare how stable each tool is, whether its highlighted passages make sense, and how it performs on known samples similar to your actual text. If all three disagree sharply after those checks, the honest result is uncertainty, not a reason to pick the detector that produced the largest number.
Watch out for treating that benchmark as settled, since @redadmin already admitted nobody confirmed who ran it. A ‘caught’ table with no false-positive column tells you which tool flags the most, not which one is right. Clever topping the humanized column is genuinely interesting, but I’d sit with the uncertainty instead of ranking anything until someone shows how these three handle clean human writing.
Feed each of them a couple pages you wrote yourself before 2022, back when nobody had ChatGPT to lean on. If a detector still calls that AI, you’ve got your answer about which score to trust on the rest. The @redadmin table tells you what each tool catches, but your own old text tells you what it wrongly accuses.
