Ask a working tech reviewer what actually changed about the job over the past two years and the honest answer usually isn’t “AI writes my reviews now.” It’s that something new sits between the reviewer and the page: a detector score. Freelancers get asked to submit a “human” percentage alongside their copy. Publications run contributor drafts through screening tools before an editor reads a word. And reviewers have started writing differently – not because a machine drafted their words, but because a machine might be asked whether one did.

That shift is measurable, not just vibe. The National Bureau of Economic Research opened a September 2025 working paper on detectors with the examples its authors had in mind: “ensuring assignments were completed by students, product reviews written by actual customers.” Product reviews are now a named use case for automated authorship detection. If you write them for a living, you are in the blast radius.
The gate moved in front of the editor
For most of the internet’s history, a bad review got filtered by a human – an editor, a community manager, or the reading public. Now the first filter is often software. Marketplaces use detectors to triage suspect customer reviews; publishers use them to screen freelancers. Some of that is defensible. A five-star “review” churned out by a bot farm is a real problem, and platforms have no other cheap way to spot it at scale.
The trouble is that the same tools get pointed at the honest reviews, because a detector cannot tell a liar from a tidy writer. It only sees patterns. And a professional product review is one of the most patterned forms of writing on the web.

Why clean, structured writing gets flagged
Think about the standard review format: what it is, the specs, how it performs, who it’s for, verdict. Sentence lengths cluster. Phrases repeat – “for the price,” “in practice,” “if you mainly shoot video.” That regularity is precisely the signal detectors are trained to catch, because machine text is regular in the same way. A reviewer who edits well is, in effect, smoothing away the little irregularities the classifier reads as human.
The research on this is fairly brutal. A 2025 ACL Findings study built a 15,000-sample dataset of genuine human writing that had been lightly polished by AI models, then ran twelve detectors against it. Minimal polishing with GPT-4o pushed detection rates to anywhere from 10% to 75%, depending on the tool. And when human text received only “extremely minor” edits from an older LLaMA model, the share flagged as AI jumped to 32.31% for ZeroGPT, 42.56% for Pangram, and 64.71% for GPTZero – the AI-polished-text false-positive study is blunt about the implication: detectors routinely flag text that a human wrote and merely tidied.
Read that again. The more a conscientious writer sharpens their own prose, the more machine-like they look. Good editing has become a false-positive vector. That is a genuinely strange incentive to hand to a profession built on clear, consistent writing.

What the error rates actually look like
The four detectors below are the ones two 2025 studies tested head to head. The pattern to notice is in the last row: the free, open-source baseline is the one most likely to accuse a human writer, and it gets there by flagging nearly everything.
| Detector | Human text wrongly flagged (FPR) | AI text missed (FNR) | Notes |
|---|---|---|---|
| Pangram | Near zero on medium-to-long passages | Near zero | The only tool that held up under a strict 0.5% false-positive cap |
| OriginalityAI | ≤3% | Up to 30–42%, worse at strict thresholds | Sensitive to “humanizer” tools |
| GPTZero | ≤3% | Up to about 50% once humanizers are applied | Stronger than OriginalityAI on false positives |
| RoBERTa (open source) | 30–78% | Looks low, but only because it flags almost everything | Both studies call it unsuitable for high-stakes use |
Figures from the NBER working paper w34223 (September 2025) and the Becker Friedman Institute’s October 2025 companion study, which tested the same four detectors on a 1,992-passage corpus that included consumer reviews. Every rate depends on the classification threshold – the knob each vendor sets, and rarely publishes alongside its accuracy headline.
The tools are not all equally broken
Here is where I part ways with the “all detectors are worthless” crowd. They are not all equally bad, and the honest position is more uncomfortable than either camp admits.
The October 2025 University of Chicago study found the commercial tools sit in a different league from the free baseline. RoBERTa misclassified 30% to 78% of human-written passages as AI, while Pangram posted near-zero error rates on medium and long text and was the only detector to stay accurate under a strict 0.5% false-positive cap – you can read the Becker Friedman Institute report and its policy-cap framework directly.
But even that good news carries qualifications the vendors don’t lead with. False positives and false negatives trade off directly: tighten a detector to catch more AI and you flag more humans. The threshold is a choice, not a fact. And the independent tests that flatter the leading tool mostly ran its previous generation, not the version being sold today. The strongest defensible claim is narrow: some commercial detectors are good enough to be worth consulting, and none is good enough to be the verdict.
How reviewers are writing around the number
Once a score stands between you and your byline, you start writing for the score. Some of that adaptation is genuinely good. Reviewers now lean harder into primary evidence: the exact test conditions, the measured numbers, photos of the unit on the bench, the one weird detail that could only come from actually using the thing for a week. Those are the traits that make text read as human to a classifier – and, not coincidentally, the traits that make a review worth reading.

The other half of the adaptation is worse. Writers sprinkle in typos on purpose. They chop sentences to lower a “burstiness” score. They run drafts through “humanizing” tools that paraphrase until the number drops, which mostly means degrading the prose until it sounds like a bored human rather than a tidy one. My read: that is the tell that something has gone wrong. The moment a reviewer optimizes for the detector instead of the reader, the reader loses – and visibly so, because a botched transition you made botched on purpose is still a botched transition.
There’s one more practical habit in the mix. Because the tool that judges your draft might be some free checker a platform runs on the back end, it’s tempting, and common, to run the draft through an AI checker free of charge yourself – just to see the number before somebody else does. Do it as a sanity check, not a verdict. A free detector that flags a clean paragraph is telling you something about the detector, not about your writing.
What I’d actually do
If I wrote reviews for a living, here is the short version of how I’d handle this:
- Keep the audit trail. Timestamped notes, raw measurement files, the original images. Not for the detector – for the editor, and for you when someone questions a claim.
- Disclose AI assistance plainly. If you used a model to check grammar or brainstorm test cases, say so. The ambiguity is worse than the use.
- Never let a score decide a writer’s fate. If you are the editor, read the submission before you run it through anything. If you are the writer, remember that a flagged draft is a starting point for a conversation, not a conviction.
- Don’t humanize. Specify. If a passage reads suspiciously smooth, the fix is a concrete number, a named condition, or an admission of uncertainty – not a deliberate typo.
The question worth arguing about
Suppose a publication’s policy is “no AI-assisted writing.” You wrote the whole review by hand. You run it through a detector, it gets flagged, and you rewrite until it passes. Did you use AI to write the piece?
I don’t think that’s cheating – it’s editing under a broken constraint. But I’d argue the constraint is the problem: a rule that forces honest writers to perform imperfection in order to prove they’re honest is not protecting anyone’s trust. And I’d be curious which way you’d call it, because plenty of working writers I’ve compared notes with land on the other side and say passing the check is just part of the job now.

FAQ
Do AI content detectors actually work?
Partially. Independent 2025 testing found the best commercial detectors can reach near-zero error rates on longer text, while free and open-source tools flag huge shares of human writing. No detector is reliable on short passages, and none should be treated as proof of authorship.
Can AI detectors tell if I used Grammarly?
Often yes, and unfairly. A 2025 study found that even extremely light AI polishing of human-written text pushed false-positive rates sharply upward. Light grammar editing can make honest writing look machine-made to some tools.
Why do AI detectors flag human writing?
Because they detect patterns, not authorship. Repetitive phrasing, uniform sentence length, and rigid structure – all hallmarks of a professional product review – look the same whether a human or a model produced them.
Should tech reviewers use AI at all?
For research, outline brainstorming, and grammar, it’s defensible if disclosed. For the actual opinions, test results, and judgments, the value of a review is the human who did the testing – that can’t be delegated without hollowing out the reason the review exists.
Does Google penalize AI-written reviews?
Google’s guidance targets unhelpful, low-effort content rather than AI authorship itself, and the company has said it does not rank or demote pages purely because of a detector score. The practical risk is low originality and weak firsthand evidence, not the detector reading.
What’s the difference between AI-generated and AI-polished text?
AI-generated text is drafted by a model. AI-polished text is human writing that a model lightly edited. Detectors struggle to separate the two, which is exactly why a writer who tidies their own prose can end up in the same bucket as someone who prompted a chatbot.
How this article was put together
This piece set out to answer one question for writers and editors who deal with product reviews: what do the primary studies on AI detection actually show, and what does that mean for how reviews get written? I read the source research rather than the vendor blogs – a 2025 NBER working paper and a University of Chicago companion study testing four detectors on a 1,992-passage corpus that included consumer reviews, plus a 2025 ACL Findings paper on AI-polished text. Where the vendors’ own numbers could not be independently checked, I said so. The detector accuracy figures will age quickly, and any of them should be rechecked before citing in 2027.






