Towards Data Science iconTowards Data ScienceSep 28, 2026 ~6 min source read

Grokking: When a Model Suddenly Learns Long After It Appears Done

A small neural network trained on modular addition memorized answers early, then—after many more training steps—abruptly acquired a general, human-like method. That late emergence, called grokking, exposes how training curves can hide deep internal change and raises practical questions for model development.

Share this story

Send the public story page.

Useful takeaways from this story.

In the modular-addition experiment, the network discovered a circular number representation and solved addition via rotation—effectively rediscovering trigonometric structure without being told.

Standard training monitoring and early stopping can cut off models before grokking appears, so flat external metrics do not guarantee stasis inside the model.

The useful part

They trained a tiny neural network, far smaller than anything that gets called "AI" today, on one of the simplest tasks imaginable: modular addition, basically clock math. Within a thousand training steps it was answering every practice question correctly. It's a familiar failure in machine learning: the model had memorized the answer key instead of learning the rule behind it.

How it works

  • Except the researchers kept training the model well past the point most people would call it finished.
  • At 1,000 steps it scored perfectly on training data but only about 10 percent on new problems.
  • By 20,000 steps, it scored close to 100 percent on both, with no change to how it was being trained in between [1].
  • Real understanding, measured by accuracy on new problems, doesn't show up until much later.
  • Stop watching them right when their quiz scores plateau, and you'd conclude they're a memorizer and move on, with no way of knowing that real understanding was still forming underneath.

What to take from it

Give them a new question that tests the same idea in a slightly different form, and they freeze, because they memorized answers rather than the concept underneath them. They stop recalling flashcards and start actually understanding the topic well enough to solve problems they've never seen. That raises a question worth taking seriously well beyond toy math problems.

Example or evidence

  • Generalization Beyond Overfitting on Small Algorithmic Datasets (2022), arXiv:2201.02177 [2] N.
  • Then the researchers tested it on questions it had never seen before.
  • Cramming versus actually understanding An analogy that captures it well: picture a student who crams the night before a test.
  • Now imagine that student keeps studying anyway, not new material, just the same material again and again.

Details worth keeping

What is 8 plus 7 if the clock only goes up to 12? They can answer yesterday's practice questions perfectly. The click happens long after the student appears finished.

Related coverage

  • E27: Throughout 2026, Claude users repeatedly reported the same practical failure: workflows that had worked reliably stopped working, long sessions lost their thread, and instruction-following became less...
  • Thenextweb: Developers can improve artificial intelligence systems in various ways.
  • Medium: This is how top developers are learning AI differently in 2026 Continue reading on Medium »
  • Generativeai: The study that everyone reported wrong has a much more specific finding buried inside it. Continue reading on Generative AI »
  • Medium: On day eleven of a real 2​026 experimenвЃ t, an AвЃ I‌ agent cast the deciding vote toвЃ delete herself, then wrote‍ in her d‍iar​y that it⁠… Continue reading on Data Science Collective В»

More context around this story.

The most valuable part of AI may not be the model
E27 iconE27Sep 9, 2026

The most valuable part of AI may not be the model

Throughout 2026, Claude users repeatedly reported the same practical failure: workflows that had worked reliably stopped working, long sessions lost their thread, and instruction-following became less dependable. The complaints did not arrive as a smooth decline. They came in bursts. Users would suddenly report that a

Loading more related stories...

Keep reading in the app

Open the app view to save this story, compare related coverage, and continue from the same source.

Open in app