I ran a post I wrote by hand through Originality.ai, the AI detector that markets itself as the strict, accurate one that publishers and agencies pay for. It came back 96% confident that’s AI.
Here is the part that still makes me laugh. The post it flagged is about AI detectors falsely flagging non-native writers. The tool read a human explanation of how it makes false positives, and responded by making that exact false positive, in real time, with 96% confidence. The machine proved the thesis by failing the way the thesis said it would.
This is my Originality.ai tested human copy write-up, the second in a short series where I run my own genuinely human writing through the popular detectors and show you the raw result. The first was my GPTZero test, which rated a different human post of mine 100% AI. Two tools now. Two of my human posts. Two false positives.
If you want the underlying mechanism, why these tools flag clean writing at all, that is covered in full in the post on why AI detectors flag your writing. This post is the field test on the tool that claims to be above this problem.
Why Originality.ai is the interesting one to test
GPTZero is the detector most people have heard of. Originality.ai is the one that gets used against you.
It is built for a specific buyer: content agencies, publishers, and SEO managers who hire freelance writers and want to police whether those writers used AI. It markets itself as stricter and more accurate than the free tools, and it charges for that promise. When a client runs your work through a detector without telling you, there is a real chance it is this one.
That makes it the important tool to test honestly. If the detector that sells itself as the professional-grade, accurate option flags clean human writing by a non-native author, the “we are more accurate” promise deserves a much closer look. So I tested it the only way that means anything: on writing I know for certain a human produced, because I am the human, and there was no AI in the process at any point.
What I tested and what came back
The text was a real LinkedIn post of mine. Its argument is that non-native writers get flagged as AI far more often than native speakers, that the cause is how detectors score predictability, and that non-native writers who write clean, even, careful sentences look “machine” to the tool precisely because they write carefully.
I wrote every word of it. No AI draft, no AI rewrite, no humanizer.
Originality.ai’s verdict:
👉🏼 96% confident that’s AI.
👉🏼 The tool’s own note under the score explains this means it is 96% confident the text is AI-generated, not that 96% of the text is AI. In other words, near-total confidence in a wrong conclusion.
👉🏼 The scan highlighted almost the entire post in red as AI-likely, including the sentences explaining why this false positive happens.
There is no authorship dispute to untangle here. The post is human. The detector that bills itself as the accurate one was 96% sure it was not.

The irony is the lesson
It would be easy to stop at “ha, the tool was wrong.” But the specific way it was wrong is the useful part for you.
Look at what my post actually says, in the lines the tool flagged: non-native writers tend to write clean, even, careful sentences, and that evenness is exactly what the machine reads as machine. That sentence describes the mechanism. The tool then flagged that sentence, and every sentence around it, because those sentences are clean, even, and careful. The post is a worked example of its own argument. I did not plan it that way. The detector planned it for me.
This is the thing to understand about your own writing. The traits that make the detector suspicious are not flaws. They are the exact craft skills you have worked to build: clear word choices, controlled sentences, low fuss. As Stanford professor James Zou, senior author of the landmark study on this bias, explained, detectors score “perplexity,” which tracks how sophisticated or surprising the writing is, and non-native writers naturally score lower on the lexical and syntactic measures the tools use. Lower perplexity reads as AI. Your clarity is the thing being punished.
❌ The assumption: a stricter, paid detector like Originality.ai will be accurate on human writing.
✅ The reality: “stricter” often just means more confident when it is wrong, and clean non-native copy is exactly what it is most confident about.
The stat behind the anecdote
One test is one data point, so let me put the peer-reviewed number next to it.
A 2023 Stanford study by Weixin Liang, James Zou, and colleagues ran 91 human-written TOEFL essays by non-native English students through seven major detectors. The tools were near-perfect on essays by U.S.-born students. They flagged 61.22% of the non-native essays as AI-generated, and 97% of those human essays were flagged by at least one detector. The false positive I got from Originality.ai is not bad luck. It is the documented base rate landing on my post.
So when a detector flags your human writing, you are not the strange exception. You are the majority case in the research, and the “5x more often than native speakers” framing from my flagged post is a fair summary of that 61% versus roughly 5% gap.
What this means for you
Three practical takeaways, the same spine as the GPTZero test because the lesson holds across tools.
The score is about the tool, not about you. A 96% AI rating on human writing measures the detector’s method, not your honesty or your skill. Once you hold that firmly, a scary screenshot loses its power to make you doubt your own work.
Do not “test and fix” your writing against these tools. Some non-native writers run their drafts through Originality.ai out of fear, see a high score, and then wreck their clean copy trying to lower it. They add clutter, break good sentences, and make the writing worse to satisfy a broken gauge. Do not damage real craft to move a fake number.
Keep proof of authorship instead of chasing a score. Since any tool can flag you at any time, your protection is documentation, not a lower percentage. Write in Google Docs with version history on so the evidence exists before anyone asks. The full client-facing playbook, what to say when a client emails you a detector screenshot, is in the post on being falsely accused of using AI.
What I will not do about this result
I will not run the post through a humanizer to beat Originality.ai, and I would tell you not to either.
Humanizers lower AI scores by adding the mess and irregularity that raise perplexity. For most writing, that means worse sentences: clumsier, less clear, less yours. For a non-native professional whose whole edge is clean, controlled, register-aware English, that trade runs backward. You would be spending your actual skill to satisfy a tool that is wrong about you in the first place.
The detector is a broken gauge on the roadside. You do not rebuild your engine to please it. You write well, you keep your receipts, and you let the gauge be wrong.
The honest limits of this test
I want to be precise about what this does and does not show.
This is one post, tested once, on one detector. It is not a controlled experiment, and I am not claiming Originality.ai flags everything, or that it is worse or better than GPTZero on identical text, because I ran different posts through each. What I can say is narrow and solid: a detector that markets itself as the accurate, professional-grade option rated a genuinely human post by a non-native writer as 96% AI, and the peer-reviewed research says this outcome is common, not freakish.
If anything, using two different posts makes the pattern more worrying, not less. It was not one unusually “robotic” piece of writing that tripped one tool. It was two ordinary posts, tripping two different tools, in the direction the research already predicted.
If it can happen to these two posts, it can happen to yours. Not because your writing is weak. Because the tools measure predictability, and your clarity looks predictable to a machine.
Frequently Asked Questions
Is Originality.ai accurate for detecting AI writing?
In my test, no. Originality.ai rated a 100% human-written post as 96% confident AI. It markets itself as stricter and more accurate than free detectors, but “stricter” can mean more confident when it is wrong. This matches peer-reviewed findings: a 2023 Stanford study found major detectors flagged 61.22% of human-written non-native essays as AI-generated. The tools measure predictability, not actual authorship.
Why did Originality.ai flag my human writing as AI?
Because it scores how predictable your word choices and sentence patterns are, not whether a human wrote them. Clean vocabulary, even sentences, and careful structure lower your “perplexity” score, which pushes the text toward the AI label. Non-native writers who write carefully in a second language produce exactly these traits, so their genuinely human work is flagged more often than native writing.
Is Originality.ai better than GPTZero?
Not in any way my tests could show. I ran different human posts through each: GPTZero rated one 100% AI, Originality.ai rated another 96% AI. Both produced confident false positives on human writing by a non-native author. Originality.ai charges for a promise of higher accuracy, but on the specific problem of flagging clean non-native copy, both tools failed in the same direction.
Should I run my writing through Originality.ai before sending it to a client?
No. Checking your own work out of anxiety often backfires: a false high score can push you to damage clean writing by adding clutter to lower the number. Instead of chasing a score on a flawed gauge, keep proof that you wrote the work, such as Google Docs version history, which is the evidence that actually matters if a client ever questions it.
What should I do if a client uses Originality.ai to accuse me of using AI?
Stay calm, do not apologize for something you did not do, and offer concrete evidence like your version history. Note that these detectors produce false positives often, especially on clean professional writing. The full step-by-step response, including the exact message to send and a contract clause that prevents the problem, is in the post on being falsely accused of using AI.
Does a 96% AI score mean my writing is bad?
No. It usually means your writing is clear and predictable, which are strengths, not faults. The detector cannot tell the difference between “predictable because a machine wrote it” and “predictable because a skilled human chose clear words.” The score describes the tool’s method, not the quality or origin of your work.
Where to go next
👉🏼 For the companion test on the other major detector, see my GPTZero result, where a different human post came back 100% AI.
👉🏼 For the mechanism behind why both tools do this, see why AI detectors flag your writing.
👉🏼 For the client playbook when someone sends you a detector screenshot, see what to do when you are falsely accused of using AI.
👉🏼 For the diagnostic on your own writing, the Natural English Edit is the 15-pattern checklist with prompts to run on your own copy. Free.
Originality.ai said my human post was 96% AI. It was 100% human. Your job is not to move that number. It is to write well and keep the proof.