Towards Data Science iconTowards Data ScienceSep 23, 2026 ~8 min source read

A Tiny Network Packed Five Features into Two Dimensions — and Drew a Pentagon

Reproducing Anthropic’s toy model of superposition in pure NumPy with hand-derived gradients shows how a small neural network can encode more features than its dimensionality by arranging them at angles, producing a pentagon when compressing five sparse features into two dimensions.

Share this story

Send the public story page.

Useful takeaways from this story.

Reproducing the toy model in plain NumPy required deriving the backward pass by hand because weights are tied (used for both encoding and decoding), which adds terms to the gradient.

With sparse synthetic features and a bottleneck smaller than the number of features, the trained model arranges five features around a two-dimensional circle — producing a pentagon — instead of selecting only two features to represent.

Hand-deriving gradients helps you understand how tied weights and ReLU activation interact in the loss gradient and makes the mechanics of superposition clearer than relying on automatic differentiation alone.

The useful part

The claim: a neural network can represent more features than it has dimensions to work with, by packing them in at angles to each other and tolerating a bit of interference. Ask it to compress five things into two dimensions under the right conditions, and it doesn't pick two winners and give up on the rest. I didn't have PyTorch or any autograd library available, and no internet access to install one either, so everything below is plain NumPy, and I derived the backward pass by hand.

How it works

  • Deriving the gradients yourself forces you to actually understand what the model is doing to the data, rather than trusting a.backward() call you've never had to think about.
  • If you have a math background and you've never hand-derived backprop through even a tiny network, I'd genuinely recommend it as an exercise.
  • I generate it myself in code, there's no external dataset involved, which is also how the original paper does it.
  • The problem this is trying to explain Here's the motivating puzzle, and it's a real one in interpretability research.
  • If you look inside a trained neural network hoping to find individual neurons that cleanly represent individual concepts, one neuron for "is this a dog," one for "is this red," you mostly don't find that.

What to take from it

This is called polysemanticity, and it makes interpretability much harder, because you can't just read off what a network "believes" by inspecting individual units. It's a real strategy the network uses on purpose, because it has more concepts to represent than it has neurons to represent them with, and most of those concepts are rarely active at the same time. Gradient descent found that arrangement on its own because it's genuinely the optimal packing, not because anyone told it what a pentagon was.

Example or evidence

  • Instead you find neurons that seem to respond to several unrelated things at once, a neuron that fires for both cat faces and the front ends of cars, say.
  • This packing strategy is what the paper calls superposition, and its toy model is designed to be the simplest possible setting where you can watch it happen and actually measure it.
  • The model, and the math I had to work out to train it The setup is small on purpose.
  • You have n synthetic features, each one a number between 0 and 1 that's zero most of the time (that's the sparsity) and nonzero the rest of the time.

Details worth keeping

I was reproducing a small piece of Anthropic's 2022 interpretability paper, "Toy Models of Superposition" [1], mostly because the central claim sounded implausible enough that I wanted to check it myself rather than take it on faith. It arranges all five into a perfect pentagon. That turned out to be the right kind of annoying.

Related coverage

  • Medium: Someone built a text generator whose entire model is one line: score each candidate continuation by how small gzip makes it. Continue reading on Medium »
  • Habr: В какой‑то момент я посмотрел на 13 ГБ весов LLaMA-7B и подумал: а почему, собственно, они должны занимать именно столько?
  • Activistpost: One site is near the military's main chemical an...
  • Techmeme: Julie Bort / TechCrunch: PrismML releases Bonsai 2 27B, which compresses Alibaba's Qwen3.8 27B to 5.9 GB, small enough for smartphones, while retaining 98.2% of Qwen's benchmark scores — If AI lab PrismML...
  • Medium: Most guides told me to buy a bigger card. I kept using the one I had, until I understood what its limits really meant. Continue reading on Medium »

More context around this story.

Я попробовал ужать LLaMA-7B до 1 ГБ без дообучения. Вот что произошло после 60+ экспериментов
Habr iconHabrSep 9, 2026

Я попробовал ужать LLaMA-7B до 1 ГБ без дообучения. Вот что произошло после 60+ экспериментов

В какой‑то момент я посмотрел на 13 ГБ весов LLaMA-7B и подумал: а почему, собственно, они должны занимать именно столько? Что, если не делать маленькую модель, не учить её заново и не заставлять копировать большую, а просто найти более короткую запись тех же «мозгов»? Здравый человек, вероятно, скачал бы готовую квант

PrismML releases Bonsai 2 27B, which compresses Alibaba's Qwen3.8 27B to 5.9 GB, small enough for smartphones, while retaining 98.2% of Qwen's benchmark scores (Julie Bort/TechCrun...
Techmeme iconTechmemeSep 18, 2026

PrismML releases Bonsai 2 27B, which compresses Alibaba's Qwen3.8 27B to 5.9 GB, small enough for smartphones, while retaining 98.2% of Qwen's benchmark scores (Julie Bort/TechCrun...

Julie Bort / TechCrunch : PrismML releases Bonsai 2 27B, which compresses Alibaba's Qwen3.8 27B to 5.9 GB, small enough for smartphones, while retaining 98.2% of Qwen's benchmark scores — If AI lab PrismML isn't on your radar yet, it should be — not because it's raised gobs of money (it hasn't yet …

Loading more related stories...

Keep reading in the app

Open the app view to save this story, compare related coverage, and continue from the same source.

Open in app