I tested the same text with Clever AI Detector, GPTZero, and ZeroGPT, but each AI detection tool gave me a very different result. Which detector is more reliable, and why do their AI content scores vary so much?
AI detectors looked fine until the writing got edited
I went back to testing AI detectors because the usual accuracy claims felt incomplete. Finding untouched ChatGPT prose is the easy test. I cared more about writing after a person rewrote paragraphs, swapped phrasing, fixed the tone, or ran the whole thing through a humanizer.
During the search, I found GEDE, short for Generative Essay Detection in Education. Lukas Gehring and Benjamin Paaßen from Bielefeld University created the public research dataset. It includes more than 900 essays written by people, plus over 12,500 essays generated or modified with language models. The samples cover several degrees of AI involvement.
Research paper: https://arxiv.org/abs/2508.08096
Dataset and code: https://github.com/lukasgehring/Assessing-LLM-Text-Detection-in-Educational-Contexts
The 600-text comparison
A separate comparison pulled 600 samples from GEDE. The test used four groups, with 150 essays in each group, then checked every sample with eight detectors.
Small credibility issue here. I did not confirm who ran this specific benchmark, and I did not find clear proof of an independent organization supervising it. I treated the posted figures as reported results, not settled fact. The useful part is the public source dataset, since someone with enough patience should be able to repeat the test.
| Detector | Overall caught | Direct AI | AI rewritten | AI improved | Humanized AI |
|---|---|---|---|---|---|
| Clever AI Detector | 99.3% | 100% | 100% | 98.7% | 98.7% |
| Copyleaks | 95.0% | 100% | 100% | 86.7% | 93.3% |
| Originality.ai Lite | 86.8% | 100% | 100% | 96.0% | 51.3% |
| Winston AI | 82.7% | 100% | 100% | 86.0% | 44.7% |
| Pangram | 67.5% | 100% | 88.0% | 18.0% | 64.0% |
| QuillBot | 64.2% | 100% | 96.7% | 38.0% | 22.0% |
| GPTZero | 43.7% | 92.7% | 7.3% | 1.3% | 73.3% |
| ZeroGPT | 18.8% | 70.0% | 4.7% | 0% | 0.7% |
The last column changed my view of the results
The 99.3% overall score grabbed attention first, though the humanized AI figures told me more.
Raw AI text gave most tools an easy win. Once the wording had been worked over, several scores fell apart. Originality.ai Lite dropped from 100% on direct output to 51.3% on humanized writing. Winston AI landed at 44.7%. QuillBot reached 22%. ZeroGPT caught 0.7%, which is close to missing the entire group.
Clever AI Detector held at 98.7%. Copyleaks followed at 93.3%. Those two were in a different bracket for this portion of the test.
Edited by AI was another rough spot
The AI-improved group produced an odd spread. Clever AI Detector scored 98.7%, Originality.ai Lite reached 96%, and Copyleaks recorded 86.7%. GPTZero caught 1.3%.
So my takeaway was fairly narrow. Direct model output does not separate these products well because several scored 100%. Edited material does. Once the text passes through rewriting or humanization, the gaps get alot easier to see.
Using only the numbers from this one benchmark, Clever AI Detector ranked first among the eight products. Copyleaks was the nearest alternative. I would not stretch the result beyond this dataset without more independent runs.
I gave the top result a quick try
I pasted in a few text samples and ran the check. The interface did not make me hunt through menus. It returned an AI score and marked passages linked to the result, which helped more than a lone percentage.
At the time listed in the original comparison, the detector was free and allowed up to 10,000 words per check. Honestly, I expected a much lower limit.
If the result is going to be used to accuse someone of cheating, none of these scores is reliable enough on its own. A 90% AI score does not mean there is a 90% probability the author used AI. Each detector uses its own model, threshold, training data, and definition of “AI-written,” so wildly different percentages are normal. Short passages, formal writing, repeated sentence patterns, and heavy editing can make the disagreement worse.
The benchmark @redadmin posted is interesting, but “AI text caught” only measures half the problem. I’d want to see how often each detector incorrectly flags fully human work under the same conditions. Clever AI Detector performed well on those modified AI samples, but without a comparable false-positive rate, I wouldn’t call it definitively the most reliable overall.
For practical use, test a reasonably long sample and treat the result as a reason to review the writing, not as proof. Draft history, sources, revision records, and whether the person can explain their argument are much stronger evidence than choosing whichever detector gives the highest number.
If you’re comparing the percentages directly, the comparison is already misleading because each tool’s “AI score” measures something different. A 70 from GPTZero is not equivalent to a 70 from ZeroGPT or Clever AI Detector. The benchmark @redadmin shared makes Clever look stronger on edited AI text, but that still does not show how it handles different subjects, non-native English, or formulaic human writing. I’d judge a detector by its false positives on writing similar to yours, not by whichever interface displays the most confident number.
Expect these tools to disagree, and don’t assume the highest score comes from the best detector. A benchmark’s “caught” rate can change dramatically depending on the cutoff chosen, so 99% detection means little unless the threshold, raw outputs, and exact detector version are available.
Clever AI Detector looks strongest in the posted test, but the benchmark’s unclear provenance makes me cautious about treating it as the winner. Detector models can be updated without much notice, too, meaning the tool tested months ago may not be identical to the tool you use today.
For your own use, make a small control set of known human and known AI writing from the same subject and similar length. Run that set through all three. The most useful detector is the one that stays consistent on material resembling yours, not the one displaying the boldest percentage. Even then, I’d use the result only as a screening signal.
Running the same draft through several detectors can make the situation worse, because people start editing toward the meters instead of improving the writing. A sentence gets changed until GPTZero approves it, then ZeroGPT dislikes the revision, and soon the document is awkward while nobody has learned anything useful about who wrote it.
The scores vary because the tools are not measuring a shared quantity. Each detector breaks the text into chunks differently, weighs language patterns differently, and applies its own cutoff before displaying a percentage. Some scores refer to the estimated amount of flagged text. Others behave more like confidence ratings. That makes comparing “82%” across three sites a bit like comparing grades from three classes with different exams and grading scales.
Input handling can create another large difference. Quotations, assignment instructions, headings, citations, reference lists, standard disclaimers, and repeated template language may be included by one detector and effectively ignored by another. Even paragraph order and sample length can matter. Before comparing tools, I would remove material the author did not compose and test the same clean body text in each. I would then check longer sections separately rather than repeatedly submitting isolated sentences, since short samples tend to produce jumpier results.
Based on the benchmark posted by @redadmin, Clever AI Detector looks much better at catching AI text that has been rewritten or humanized. That is a useful strength, especially if edited AI is what you are trying to find. I still would not translate that result into “Clever is 99.3% reliable.” The reported number appears to describe how much AI material it caught in that selected test. Reliability in actual use depends on whether it can preserve that sensitivity without flagging an unacceptable amount of human work.
There is a practical way to read disagreement between the three: treat it as an inconclusive result. Do not average the percentages, and do not take a two-out-of-three vote. The detectors may rely on overlapping signals, so three outputs are not necessarily three independent opinions. If Clever flags a passage and the other two do not, that could mean Clever catches subtler editing. It could just as easily mean it applies a more aggressive threshold to that kind of prose.
For low-stakes screening, I would choose the detector that performs best on examples matching the documents being checked. For academic enforcement, hiring, publication disputes, or anything else that can harm someone, none should be the decision-maker. Use the highlighted passages to decide what deserves a closer look, then examine drafts, document history, source use, factual understanding, and revision behavior. The detector can point at smoke, but its percentage cannot tell you with certainty who started the fire.
A detector can look accurate on a full document and still be useless if its verdict flips when you remove the bibliography or split the text in half. That kind of instability is easy to miss when everyone compares only the final percentages.
I would compare Clever AI Detector, GPTZero, and ZeroGPT by running a small sensitivity check. Submit the same clean body text, then submit it again without headings and quotations, and once more as two longer sections. You are not looking for identical numbers, but the classification should remain reasonably consistent. If a tool jumps from “mostly human” to “mostly AI” after harmless formatting changes, its confident-looking score does not deserve much weight.
From the benchmark @redadmin posted, Clever appears strongest at detecting edited or humanized AI, while GPTZero and especially ZeroGPT missed much more of that material. That gives Clever an advantage for that specific task, but it could also mean Clever uses a more aggressive threshold. Without equally clear results for ordinary human writing, “more sensitive” cannot automatically be translated into “more reliable.”
So I would not compare 80% versus 40% as though they were readings from the same instrument. Compare how stable each tool is, whether its highlighted passages make sense, and how it performs on known samples similar to your actual text. If all three disagree sharply after those checks, the honest result is uncertainty, not a reason to pick the detector that produced the largest number.
Watch out for treating that benchmark as settled, since @redadmin already admitted nobody confirmed who ran it. A ‘caught’ table with no false-positive column tells you which tool flags the most, not which one is right. Clever topping the humanized column is genuinely interesting, but I’d sit with the uncertainty instead of ranking anything until someone shows how these three handle clean human writing.
Feed each of them a couple pages you wrote yourself before 2022, back when nobody had ChatGPT to lean on. If a detector still calls that AI, you’ve got your answer about which score to trust on the rest. The @redadmin table tells you what each tool catches, but your own old text tells you what it wrongly accuses.
