Towards Data Science iconTowards Data ScienceSep 26, 2026 ~7 min source read

AI Slop Is Already in Your Training Dataset — Three Cheap Checks, One Unexpected Result

A hands-on test on a 2019 IMDb movie-review collection found that common heuristics for spotting AI-written text flag many genuine reviews and that filtering by those flags can make a downstream sentiment model less accurate.

Share this story

Send the public story page.

Useful takeaways from this story.

At a setting that captured 80% of generated reviews, embedding density also labeled 47% of the 2019 IMDb reference reviews as possibly AI written.

Before filtering training data for model-generated text, measure how detection and removal affect the target model’s performance.

The useful part

They applied the method to peer reviews submitted to four major artificial intelligence conferences after ChatGPT launched. They estimated that 6.5 percent to 16.9 percent of the review text showed signs of substantial AI modification. These were peer reviews written by researchers making careful technical judgments in settings with real submission consequences.

How it works

  • It shows that substantially AI modified writing appeared in consequential peer review, not how common it is in product reviews, forums, or surveys.
  • For years, the standard complaint about data science work has been that most of it is cleaning messy data.
  • Most data cleaning checks catch missing values, repeated entries, and fields outside an expected range.
  • I tested three checks on the 600 reviews, measured how many generated and source reviews they flagged, and measured how filtering affected the sentiment model.
  • At a setting that caught 80 percent of generated reviews, embedding density, which measures how closely a review resembles others, also flagged 47 percent of the source reviews.

What to take from it

A newer problem is that the mess can include fluent text written by a model to sound human. I treated the IMDb reviews as human-written references because the collection identifies IMDb as their source and was published in 2019, before ChatGPT's public release. AI writing can receive a lower perplexity score because it often uses familiar word patterns, but human writing can be predictable too.

Example or evidence

  • I ran two tests using a movie review collection published by Mendeley in 2019.
  • For the source reviews, I used 1000 Movie Reviews for Reputation Generation, published by Abdessamad Benlahbib on Mendeley Data on 9 March 2019.
  • The collection contains 1,000 reviews of 10 films, each with a manually assigned positive or negative sentiment label.
  • The sample included 400 generated reviews with known origins.

Details worth keeping

That estimate is specific to conference reviews. A separate 2024 Nature paper, AI Models Collapse. The study documents this failure mode under repeated training on generated text.

Related coverage

  • Qualitydigest: Your AI Model Is Drifting Right Now. Would You Know? A validated model went quietly wrong. Here's the ladder that catches it. Quality Digest Tue, 09/22/2026 - 12:03 Off Sachin Bhandari Off
  • Medium: Why a perfectly clean dataset can still produce a dangerously misleading machine learning model Continue reading on Medium »
  • Qualitymag: Conventional screening asks: does this unit violate a limit we set?

More context around this story.

Loading more related stories...

Keep reading in the app

Open the app view to save this story, compare related coverage, and continue from the same source.

Open in app