Github iconGithubAug 27, 2026 ~1 min source read

The Load-Bearing Vocabulary of Claude

GitHub pull request descriptions, grouped by the words they are written with rather than by anything they were told to look for: eight ways of writing, and every description belongs to one of them. One of the eight was 1.0% of the corpus at the start of 2025 and is 45% of it by the middle of 2026.

Share this story

Send the public story page.

Useful takeaways from this story.

GitHub pull request descriptions, grouped by the words they are written with rather than by anything they were told to look for: eight ways of writing, and every description belongs to one of them.

One of the eight was 1.0% of the corpus at the start of 2025 and is 45% of it by the middle of 2026.

That means different answers to the same prompt become more similar to each other — there's less variety to watermarked text responses.

Building the complete brief

The page is ready to read now. The fuller skim-friendly version will appear here automatically.

The useful part

GitHub pull request descriptions, grouped by the words they are written with rather than by anything they were told to look for: eight ways of writing, and every description belongs to one of them. One of the eight was 1.0% of the corpus at the start of 2025 and is 45% of it by the middle of 2026. That means different answers to the same prompt become more similar to each other — there's less variety to watermarked text responses.

How it works

  • If you want to get a little more depressed, I've been doing deeper reading into AI watermarking research, and one of the admissions in Google's white paper on SynthID-Text — the watermarking scheme...
  • That's a load-bearing problem given that Claude clearly already struggles in this regard before watermarking its output.

Keep reading in the app

Open the app view to save this story, compare related coverage, and continue from the same source.

Open in app