A Client Said My Writing Looks AI-Generated. Here’s What I Did

A client ran my article through GPTZero and got 85%. I wrote every word. Here is the actual WhatsApp exchange, the second detector that said something completely different, and the one move that resolved it in under two hours.

At 8:22 PM, a client messaged me a screenshot. She’d run an article I wrote through GPTZero, and it had come back flagged. Her message: The detector platform was 85% confident the content was AI written.

I wrote every word of that article myself.

Client questioning AI detector score and Offering Google Docs version history as proof

What follows is the actual exchange, because when a client says AI generated content is in your draft, what you do in the next twenty minutes decides whether you keep the relationship. I got lucky in one respect: I had the receipts. That’s not luck you should rely on, so I’ll show you how to build it.

What the detector actually said, versus what she read

Here’s the first thing worth slowing down on, and I missed it in the moment too.

The client’s message said the detector was 85% confident it was AI written. But look at what GPTZero actually returned: AI 15%, Mixed 85%, Human 0%.

GPTZero showing Mixed 85 percent on human-written article

The tool’s own verdict line said it was highly confident the text was a mix of AI and human. Underneath, it flagged 65 of 66 sentences as likely AI generated.

“Mixed 85%” and “85% chance you used AI” are not the same claim. One is a category label, and the other is a probability of guilt. The client compressed the first into the second, which any reasonable non-specialist would do, because the interface practically invites it.

That gap matters. Half of what makes an AI accusation feel unanswerable is that the number sounds like a measurement. It isn’t. It’s a classification with a confidence attached, and the label it landed on was “mixed,” not “AI.”

Why arguing doesn’t work

My first reply was the honest one: I understood why it looked concerning, but I wrote it myself.

That did nothing. She came back with the reasonable objection, which I want you to sit with because it’s the one you’ll always face: she knew the tools weren’t always accurate, but 85% is high, and she couldn’t send work to her client if they’d run the same check and get the same result.

Client questioning AI detector score and Offering Google Docs version history as proof

Notice her real problem. It was never “Did Imtiaj cheat?” It was “I’m exposed downstream.” She wasn’t accusing me. She was worried about her own risk. Almost every AI accusation is this, and if you answer the accusation instead of the risk, you lose.

My word against a number is a fight I can’t win. She has a screenshot with a percentage. I have an assertion. Hers looks objective, and mine looks like what a guilty person would also say. Any freelancer who has been here knows the sinking feeling: everything you say next sounds exactly like a denial.

So stop denying, and produce something the detector can’t.

The system that worked: show the process

I told her I’d written it in a Google Doc, and I could send her the version history. She’d be able to watch me build it section by section, delete things, and rewrite parts. I also still had my research notes from before the draft existed.

Client questioning AI detector score and Offering Google Docs version history as proof

She asked for the history. I shared the doc.

This is the whole system, and I want to name it The Process Defense. You cannot argue with a score, but you can show the work that made the draft. A detector analyzes a finished text. Version history shows the text becoming. AI output arrives whole; human drafts arrive in messy layers, with abandoned sentences and reordered paragraphs and a paragraph you clearly hated and rewrote four times.

That’s evidence a score cannot produce and a generated draft cannot fake.

Her reply after opening it: she could see the edits in the document. Then she apologized, and explained she had to be careful because some of her clients are picky about AI.

The second detector, which made my point stronger

I added one more thing before she checked the doc. I suggested she run the same article through a different detector, and told her she’d probably get a completely different number.

Client reporting Originality.ai gave 57 percent on the same article

She tried Originality.ai. It said 57%.

Same article. Same words. Two tools, 85% and 57%, and neither of them was right, because the true figure was zero. She sent a laughing emoji, and that was the moment the score stopped being evidence in her mind. Not because I argued it down, but because she watched two authorities contradict each other with her own hands.

If you take one tactic from this post, take that one. Don’t tell a client detectors are unreliable. Ask them to run a second one. The disagreement does the arguing for you, and it costs you nothing because you already know what’s coming. I’ve written up what happens when you run several tools on identical text in the three-detector roundup, and the disagreement is the norm, not the exception.

What I said at the end, and why it mattered most

She apologized. I told her no worries, that I completely understood, and that I’d probably ask the same thing in her position. Then I made an offer:

Client accepting the version history evidence

Anytime something gets flagged, just send it to me, and I’m happy to show you the drafts and the history.

That last line is the actual deliverable of this whole episode. It converts a one-time confrontation into a standing procedure. Now nobody has to accuse anybody. She has a defined thing to do when a score spooks her, and I have a defined thing to hand over. The next flag, if it comes, is admin instead of conflict.

Here’s the register difference that matters when you write that message:

❌ I already told you I wrote it myself. These detectors are unreliable, and you shouldn’t trust them.
✅ Totally fair, I’d ask the same in your position. I’ve got the version history, want me to send it?

The first is correct and loses the client. The second concedes nothing about your work and gives her a way out of her worry. You’re not apologizing for something you didn’t do. You’re solving the problem she actually has.

Build the receipts before you need them

I could do all of this because the evidence already existed. That was habit, not foresight, and it’s the part you can copy today.

Write in a tool that keeps version history. Google Docs is the obvious one, and its history view is legible to non-technical clients. Keep your research notes and outlines in the same folder as the draft, and don’t delete them when you deliver. If you use AI anywhere in your process, for research, outlining, or a naturalness pass, keep that separate and be able to describe exactly what it touched, because a defense that’s partly false collapses entirely.

None of this is about proving you’re pure. It’s about being able to answer “show me” instead of “trust me.” For the broader script covering harder situations, including clients who won’t accept the process evidence, see the AI accusation playbook.

The accusation landed at 8:22 PM. It was fully resolved by 10:06. Not because I defended myself well, but because I had something better than a defense.

Frequently Asked Questions

What should I do when a client says my writing is AI generated?
Don’t argue with the score. Offer process evidence instead: version history from Google Docs or a similar tool, your research notes, and your outline. A detector analyzes finished text, while version history shows the draft being built in layers, which generated text cannot replicate. Ask the client to run a second detector too, since the tools frequently disagree.

Can Google Docs version history prove I wrote something myself?
It is the strongest practical evidence most writers have. It shows the document forming over time with deletions, rewrites, and reordering, which looks fundamentally different from text that arrived complete. It is not legal proof, but for a client relationship it is usually decisive.

Why do two AI detectors give completely different scores for the same text?
Because they use different models, thresholds, and training data, and they are estimating writing style rather than measuring authorship. In the exchange above, GPTZero returned 85% mixed and Originality.ai returned 57% on identical text that contained no AI writing at all. Asking a client to run a second tool often resolves the issue faster than any argument.

Does a high AI detector score mean the writing is bad?
No. Detectors flag predictability, not quality. Clean, well-structured, formal writing often scores higher precisely because it is consistent, which is why careful writers and non-native writers get flagged disproportionately.

How do I stop this happening repeatedly with the same client?
Make a standing offer: any time something gets flagged, they send it over and you show the drafts and history. This turns an accusation into a routine procedure and removes the emotional charge from future flags.

Where to go next

👉🏼 For the same tool flagging my own hand-written copy at 100%, see the GPTZero field test.

👉🏼 For the data behind false positives, especially if you write in English as a second language, see the false positive problem.

👉🏼 For why clean, careful writing scores highest, see why AI detectors flag academic writing.

👉🏼 For the full set of patterns that mark non-native copy, see the three mistakes that mark non-native copy.

You cannot win an argument against a number. You can make the number irrelevant by showing the work behind it.

Imtiaj Choudhury

Imtiaj Choudhury

Imtiaj Choudhury — non-native English copywriter in Shenzhen. Engineer turned writer, I write product pages, campaigns, and video scripts for global tech brands in English, my second language. This blog breaks down the process: how to write naturally, use AI well, and build a writing career regardless of where you're from. Father, photographer, and very slow gardener.

Articles: 45