Breaking

Sunday, September 6, 2026

My Adventure in AI-Assisted "Writing" and Detection


My Adventure in AI-Assisted "Writing" and Detection
Man vs. Machine: My Battle With AI Detectors




Where It All Started

A few months ago, I had a deadline and not much patience.

I opened ChatGPT, typed a rough outline, and let it write a full draft in about ninety seconds. It read fine. Clean sentences, correct grammar, decent structure.

But something felt off. Not wrong, exactly. Just... generic.

That small feeling turned into a genuine rabbit hole. I wanted to know: could anyone actually tell this was AI-written? And if they could, how?

The First Draft Problem

My first attempt at "AI-assisted" writing wasn't really assisted. It was just AI, wearing my name.

I copy-pasted the output, changed a sentence or two, and called it done. It was fast. It was also hollow. Reading it back, I couldn't find a single sentence that sounded like me.

That's when I started wondering if a machine could tell the difference too.

Enter the AI Detectors

I ran that first draft through a few detection tools out of curiosity.

The results were humbling. One tool flagged it as 98% AI-generated. Another gave it a "mixed" score. A third said it was probably human.

Three tools, three different verdicts, same exact paragraph. That inconsistency became the actual theme of this whole adventure.

GPTZero

GPTZero was one of the first AI detectors to get mainstream attention, originally built by a student trying to help teachers spot AI-written essays.

It uses two main signals: perplexity, which measures how predictable the word choices are, and burstiness, which looks at how much sentence length and rhythm vary. Human writing tends to be messier and more uneven. AI writing tends to be smoother.

My ChatGPT draft scored high on "likely AI." Smooth, predictable, evenly paced. The machine had, in a sense, caught its own kind.

Originality.AI

Next I tried Originality.AI, which is popular with SEO agencies and publishers who need to screen large volumes of content quickly.

It combines AI detection with a plagiarism checker, which turned out to be useful. It didn't just tell me the text looked machine-generated, it also flagged a couple of phrases that were suspiciously close to text already published elsewhere. Turns out even AI models have their favorite phrasing.

Winston AI

Then there was Winston AI, which several independent reviewers rank as one of the more accurate detectors available in 2026, especially for longer-form content.

This one felt the most thorough. It highlighted specific sentences rather than just giving a single score, which made it much easier to see exactly where the "AI-ness" was concentrated. Unsurprisingly, it was the paragraphs I hadn't touched at all.

Turnitin

Because I'm nosy, I also ran a sample through Turnitin, the tool most students know from a very different context: catching plagiarism in school.

Turnitin added AI detection features a while back, and it's now used by universities everywhere to screen student essays. Running my own writing through it felt strangely like being sent to the principal's office as an adult.

The Rewrite Experiment

Once I knew the draft would get flagged, I wanted to see if I could fix that without starting over.

So I rewrote it. Properly this time. I kept the AI's structure and rough logic, but rewrote every sentence in my own voice, added a personal anecdote, and cut anything that sounded like a textbook.

I ran the new version through the same three detectors.

The results improved a lot, but not completely. One tool called it "likely human." Another still flagged a couple of paragraphs I'd barely touched. The truth was uncomfortable: detection wasn't really measuring dishonesty, it was measuring style.

Why Detectors Disagree So Much

This is the part that surprised me most.

AI detectors aren't oracles. They're statistical guesses based on patterns learned from other AI-generated text. When a model writes in a very "average" way, predictable words, even sentence lengths, common structures, it trips the alarm. When a human happens to write in that same predictable way, they can trip the alarm too.

That's why non-native English speakers and people who simply write in a plain, structured style sometimes get falsely flagged as AI. Several of the tools I tested even include their own disclaimers admitting that no detector is 100% accurate, and that scores should be treated as a signal rather than a final verdict.

I found this reassuring and unsettling in equal measure.

The "Humanizer" Trap

Curious how far this cat-and-mouse game goes, I tried something a little sneaky: running my AI draft through a "humanizer" tool designed to make AI writing evade detection.

It worked, sort of. The detection scores dropped. But the writing quality dropped with it. Sentences got slightly awkward, word choices got odd, and a couple of phrases stopped making complete sense.

It felt like watching someone try to disguise a car as a bicycle. Technically possible. Not actually convincing up close.

Detection companies know this trick exists too. Several tools I read about are now specifically trained to catch "humanized" or paraphrased AI text, which means the arms race is very much still running in both directions.

What Actually Fooled the Detectors

Here's the honest answer: nothing exotic fooled them consistently. What worked was just... writing more like a person.

Specific details helped. A real memory, a specific number, an opinion with an edge to it. Varying my sentence length on purpose, some short, some long, some almost too long, helped more than any trick I tried.

Ironically, the thing that beat AI detection best was doing what good writers were always told to do anyway: be specific, be uneven, be yourself.

What This Taught Me About My Own Writing

Somewhere in this process, the experiment stopped being about beating a detector and started being about noticing my own habits.

I write in fairly even sentences by default. I default to safe transitions. I lean on the same handful of structures. In other words, some of my "human" writing wasn't that far from AI writing to begin with.

That was a slightly humbling discovery. Maybe the real value of this whole adventure wasn't learning to dodge detection tools. It was learning to notice when my own writing had gone flat, whether a machine wrote it or I did.

Where AI Assistance Actually Helped

To be fair to the technology, AI wasn't useless in this process. It was genuinely helpful for a few specific things.

  • Getting past a blank page when I had no idea where to start
  • Suggesting a structure I could argue with and improve
  • Catching repetitive phrasing in my own drafts
  • Offering alternative phrasings I could reject or reshape

Where it consistently failed was voice. It couldn't replicate my specific way of being wrong, my specific jokes, or the small personal details that make writing feel like it came from somewhere real.

The Bigger Question: Does Detection Even Matter?

The longer I spent on this, the more I started questioning the premise itself.

Detection tools were built for a world where "AI-written" and "human-written" were two clean, separate categories. In practice, almost everything I now write is somewhere in between: a human idea, an AI-assisted structure, a human rewrite, maybe an AI-suggested edit at the end.

That blend is becoming the normal way people write, not the exception. Publishers, teachers, and platforms are still catching up to that reality, which is partly why detection scores can feel so inconsistent and, at times, unfair.

What I'd Tell Someone Starting the Same Adventure

If you're about to go down this same road, a few honest notes from someone who already has.

Don't trust a single detector's score as the final word. Run the same text through two or three tools before drawing any conclusions.

Don't try to trick detectors with humanizer tools. It rarely holds up, and it usually makes the writing worse in the process.

Do use AI for structure and momentum, but rewrite the actual sentences yourself. That's where your voice lives, and it's the one thing a detector, and a reader, can actually feel.

Final Thoughts

I started this adventure trying to answer a simple question: can a machine tell if another machine wrote something?

The real answer turned out to be more interesting than yes or no. Sometimes it can. Sometimes it guesses wrong. And somewhere in between, I ended up learning more about my own writing than about any algorithm.

That, honestly, was the part worth the deadline stress.

No comments:

Post a Comment