Trust & Safety

Detecting AI-generated reviews in 2026: what still works, what does not

Every review platform now faces the same problem: LLMs write plausible fake reviews at industrial scale. Here is what the moderation stack actually catches — and what it misses.

ScoreReview UK Newsroom Published 19 February 2026 10 min read Trust
Detecting AI-generated reviews in 2026: what still works, what does not
Trust — ScoreReview UK Newsroom
Discuss

The cost of generating a convincing fake review has dropped by two orders of magnitude since early 2023. A broker can now produce 10,000 grammatically flawless, sector-specific reviews for the price of a small takeaway. Detection has had to change fundamentally: language-only classifiers no longer work in isolation. This article explains the current moderation stack, the categories of signal that still hold, and the honest limits of what any platform can do.

Why language classifiers alone have failed

In 2022, a classifier trained on ChatGPT output caught 88% of fake reviews at a 3% false-positive rate. The same classifier tested against 2025 model output catches under 41%. Language is no longer a distinguishing feature — modern LLMs write more fluently than most real reviewers.

What still works: provenance

The single strongest signal is whether a review is tied to a real transaction. Invitation tokens, receipt uploads, and identity verification cannot be forged by an LLM. This is why ScoreReview UK's verification ladder is the load-bearing part of the moderation stack, not a marketing tier.

What still works: reviewer graph analysis

LLMs write reviews. They do not build years of coherent, cross-platform reviewer history. Graph analysis — who else has this account reviewed, when did they join, what other identities do they share IP ranges with — remains a strong signal because it operates on behaviour, not text.

  • Account age and prior review distribution
  • IP / device / geolocation coherence
  • Timing patterns across the reviewer's history
  • Cross-platform handle correlation

What still works: transactional grounding

A real customer names the plumber, mentions the invoice number, references the day of the week. LLMs generate plausible reviews but avoid specifics they cannot verify. Reviews with zero verifiable transactional detail cluster hard in the suspect distribution.

What no longer works: perplexity and burstiness

The 2023-era detectors that measured statistical text patterns give near-random results on 2025 model output. Any vendor still selling perplexity-based detection as a primary defence is selling a 2023 product.

Human moderation is not optional

Automated systems handle roughly 92% of decisions with high confidence. The remaining 8% require a human moderator with sector context. Platforms without a human review layer publish more fakes — the maths is unforgiving.

What the regulator expects

The CMA's 2025 review-platform guidance treats 'reasonable steps' as a stack of controls, not a single detector. A platform that relies on any one signal — however clever — is undershooting.

FAQ

Can a determined attacker still get fakes through?

Occasionally, yes. No platform can promise zero fakes. What a good platform can promise is a moderation stack, a public transparency log, and a fast dispute path when a fake slips through.

Will identity verification become mandatory?

For high-risk sectors under active enforcement, yes — it already is at Tier 4. For low-risk sectors it remains opt-in, because friction reduces genuine submissions faster than it reduces fake ones.

Keep reading

Trending#DMCC#CMA#compliance#UK law#moderation#GDPR#data protection#retention#AI#reply automation#brand voice#sentiment#analytics#hospitality

Discussion (2)

Comments are stored locally on your device for this demo. Be respectful — no spam, no personal attacks.

  • Priya S.· 2 days ago

    Really practical breakdown — the four-part reply structure is now on our till-side crib sheet. Thank you.

  • Dan (Cannock Plumbing)· 5 days ago

    Went from 12 reviews to 47 in three months following almost exactly this playbook. It works.