GPTZero vs Originality.ai vs Copyleaks: Tested With Real ESL Writing

Three detectors, three of my hand-written LinkedIn posts, three confident verdicts of AI. GPTZero said 100%. Originality.ai said 96%. Copyleaks said 100%, with zero human words. The full comparison, and what it settles.

Three AI detectors tested against my own writing. Three posts I wrote by hand, with no AI draft and no AI rewrite. Three verdicts.

GPTZero: 100% AI. Originality.ai: 96% confident it’s AI. Copyleaks: 100% AI, with a word count that stings. It scored 224 words as AI text and exactly zero words as human. According to the market’s three most-used detectors, I do not exist.

This post puts the three results side by side, shows what each scan actually looked like, and answers the question the series has been building toward: do any of these tools handle real non-native writing correctly? I ran the first two tests in the GPTZero post and the Originality.ai post. This one adds Copyleaks and closes the loop.

The scoreboard

DetectorWhat I testedVerdict on my human writing
GPTZero (Model 4.6b)A LinkedIn post about register100% AI, “highly confident”
Originality.ai (Lite 1.0.2)A LinkedIn post about detector bias96% confident that’s AI
CopyleaksA LinkedIn post about AI disclosure100% AI. 224 AI words, 0 human words

Three different posts, all mine, all written the way I always write. Not one tool got close. The best any of them managed was Originality.ai leaving itself 4% of doubt.

One honest note on method before we go on. This is not a controlled experiment where the same text runs through all three tools. Each tool got a different post because the tests happened weeks apart as this series developed. If anything, that makes the result harder to dismiss. It was not one unusual piece of writing tripping one tool. It was three ordinary posts, tripping three different tools, exactly the way the research predicts for non-native writers.

The Copyleaks test, in detail

The new data point is Copyleaks, so here is what the scan showed.

The post I ran through it was about a question most freelance writers are quietly wrestling with: when a client asks whether you use AI, do you tell them? It lays out the logic for honesty and the logic for staying quiet. I wrote it by hand, the same way I wrote the other two.

Copyleaks returned a 100% AI Content Found score. It counted 224 words of AI text and zero words of human text. The sensitivity level was set at 2 of 3, the default, not the paranoid maximum. The scan also flagged “5 AI Phrases,” meaning five phrases that appear often in AI-written text. Highlighted in that pink wash: my opening question, my transitions, my careful parallel structure. The craft, in other words.

Copyleaks screenshot showing a human-written LinkedIn post rated 100% AI with 224 AI words and 0 human words

There is a small irony here worth naming, quieter than the Originality.ai one but real. The post Copyleaks flagged is about whether to disclose AI use to clients. A writer who never just copy-pastes AI texts, asking out loud how to talk about AI honestly, gets told by the machine that the asking itself was machine-made.

One more thing from that scan. Copyleaks advertises a paid feature called AI Source Match, which claims to find AI content matched elsewhere on the web. I could not verify what, if anything, it matched, because the feature sits behind a paywall. So I will not claim it indexed my post. But the design raises a fair question: if a false positive can be stored as “AI text seen elsewhere,” the error can compound. A tool that mislabels your human writing once may treat the label as evidence later.

Why three different tools failed the same way

The short version, because the mechanism post covers it in full: all three tools score predictability, not authorship. Clear words, even sentences, and controlled structure read as low perplexity, and low perplexity reads as machine. Non-native professionals write clean and careful on purpose, because careful feels safe in a second language. So the better your professional English gets, the more machine-like you look to a statistics engine.

The peer-reviewed number behind all three of my results comes from the Stanford study by Liang, Zou, and colleagues: seven major detectors flagged 61.3% of human-written essays by non-native English speakers as AI, against roughly 5% for native-born writers. My three-for-three is the extreme end of a documented pattern, not a fluke.

And it is not only non-native writing that exposes these tools. Kinja’s 2026 roundup of independent detector testing reports that in Supwriter’s 150-sample test, not a single major tool exceeded 80% overall accuracy: Originality.ai came in at 79%, Copyleaks at 77%, GPTZero at 76%. Every one of them advertises 95% or better. The same roundup notes Copyleaks misclassified about 1 in 20 human documents in benchmark testing, and that it blocked one competitor’s evaluation team from benchmarking it at all. The gap between the marketing and the measured performance is the whole story of this category.

So which detector “wins”?

None of them, for the question that matters to you.

If the question is “which tool correctly recognizes real human writing by a non-native professional,” the answer from my three tests is: none of the three. GPTZero and Copyleaks were maximally wrong at 100%. Originality.ai was nearly maximally wrong at 96%. Choosing between them is choosing which confident error you prefer.

If the question is “which tool will a client most likely use on me,” the practical answer: Originality.ai and Copyleaks in professional content work, GPTZero everywhere else. Which means the protection is the same regardless of tool, and it is not a better score.

❌ Trying to write in a way that passes GPTZero, then Originality.ai, then Copyleaks, then whatever comes next.

✅ Writing well once, and keeping version history that proves authorship against any tool, forever.

The full documentation system and the exact client conversation are in the false accusation playbook. The decision framework for how much attention this whole issue deserves is in the 2026 verdict post. If you want to check what your own drafts look like before a client ever sees them, the Natural English Edit is the 15-pattern checklist I use on my own copy. Free.

The honest limits

Three tests, three tools, one writer. That is a field report, not a study. A different writer, a different topic, or a longer sample might score differently, and each of these tools updates its models often enough that the exact numbers will drift.

What does not drift is the direction. The independent benchmarks, the Stanford data, and now three separate hands-on tests all point the same way: these tools are unreliable on exactly the kind of writing non-native professionals produce, and their confidence does not track their accuracy. A tool that is wrong at 96% confidence is more dangerous than a tool that is wrong at 55%, because the number persuades people.

You cannot control which detector a client runs. You can control whether the verdict has any power over you. Receipts beat scores. That has been true across all six posts in this series, and it is the place this comparison lands too.

Frequently Asked Questions

Which AI detector is most accurate: GPTZero, Originality.ai, or Copyleaks?
None of the three reliably recognized real human writing by a non-native professional in my tests. GPTZero rated one human post 100% AI, Originality.ai rated another 96% AI, and Copyleaks rated a third 100% AI with zero human words. Independent testing reported by Kinja found none of the major tools exceeded 80% real-world accuracy, despite marketing claims of 95% or higher.

Why did Copyleaks flag my human writing as 100% AI?
Copyleaks, like GPTZero and Originality.ai, scores how predictable your word choices and sentence patterns are, not who wrote the text. Clean vocabulary, even sentence lengths, and careful structure read as low perplexity, which the tool labels AI. Non-native writers who write carefully in a second language produce these traits naturally, so their genuinely human work gets flagged at much higher rates.

What does Copyleaks AI Source Match mean?
It is a paid feature that claims to match AI-generated content the tool has seen elsewhere on the web. A free scan does not show what was matched, so a locked AI Source Match panel is not evidence about your specific text. The related “AI Phrases” count flags phrases that appear frequently in AI-written text, which overlap heavily with ordinary professional phrasing.

Do AI detectors flag non-native English writers more often?
Yes, and it is documented. The Stanford study by Liang and Zou found seven major detectors flagged 61.3% of human essays by non-native English speakers as AI, compared with roughly 5% for native-born writers. The cause is mechanical: simpler vocabulary and more regular structure lower perplexity scores, which detectors misread as a machine signal.

Should I test my writing against all three detectors before submitting to a client?
No. Chasing a passing score across three tools with three different models is a moving target, and editing your writing to please them usually makes it worse. The durable protection is process evidence: write in Google Docs with version history on, keep your notes and drafts, and use that trail if a client ever questions authorship.

Is there any AI detector that is safe for clients to use on freelance work?
Not as a standalone verdict. Independent benchmarks put every major tool’s real-world accuracy between roughly 62% and 88%, with false positive rates that hit non-native writers hardest. Even the detector companies say scores are probability estimates, not proof. A client who treats a score as a verdict is misusing the tool, and the professional response is documented authorship, not a lower number.

Where to go next

👉🏼 For the mechanism all three tools share, see why AI detectors flag your writing.

👉🏼 For the individual field tests, see the GPTZero result and the Originality.ai result.

👉🏼 For the client conversation when any of these tools is used against you, see the false accusation playbook.

👉🏼 For how much worry this actually deserves, see the 2026 verdict.

👉🏼 For the diagnostic on your own drafts, the Natural English Edit is the 15-pattern checklist with prompts to run on your own copy. Free.

Three tools. Three human posts. Zero correct verdicts. Keep the receipts and keep writing.

Imtiaj Choudhury

Imtiaj Choudhury

Imtiaj Choudhury — non-native English copywriter in Shenzhen. Engineer turned writer, I write product pages, campaigns, and video scripts for global tech brands in English, my second language. This blog breaks down the process: how to write naturally, use AI well, and build a writing career regardless of where you're from. Father, photographer, and very slow gardener.

Articles: 45