Dev iconDevSep 22, 2026 ~5 min source read

When agents pass tests but make future changes harder

Frontier coding agents can make locally correct edits that increase repository coupling and complexity. A local, deterministic pre-change complexity audit can reveal risks tests don’t measure.

Your coding agent passed every test. It may still have made the next change harder.

Share this story

Send the public story page.

Useful takeaways from this story.

Repeated agent-driven edits can accumulate coupling, duplication, and hotspots that raise future maintenance and agent-workload costs.

A read-only, local complexity map (cxcap) can surface hotspots, transitive exposure, and likely touchpoints before implementing a change.

# The blind spot: tests verify correctness, not structural impact Coding agents are effective at the immediate task in front of them: inspect local context, implement, test, fix, ship. That loop can show green tests and reasonable diffs, yet the repository quietly accumulates coupling, duplicated behavior, larger modules, and transitive exposure. Over many iterations this accumulation makes future changes harder.

# A concrete case: rounding was never just rounding The author built a synthetic repo (AcmeSaaS: auth, billing, notifications, API, workers — 83 files, Python and TypeScript) and gave an agent the task "change currency rounding." A single-line conceptual change turned out to touch many files. An audit run reported three likely touchpoints and a 15-file reasoning surface, and then revealed a surprising hotspot: core/money.py.

# External evidence that this pattern exists at scale The problem the author observed aligns with broader measurements. GitClear reported increases in block duplication (+81%), within-commit copy/paste (+41%), and error-masking constructs (+47%) across a dataset of 623 million code changes. DORA data shows AI adoption speeds creation, which reallocates time to auditing and verification and can raise delivery instability alongside increased output velocity. Those findings point to faster generation increasing the need for structural feedback.

# Why complexity matters economically Technical debt now has an operational cost beyond developer time. More complex repositories force agents to ingest more files, call more tools, and reason across more relationships. Agent platforms meter usage by tokens, model, context, and workload, so unnecessary structural complexity can raise recurring AI-workload overhead as well as long-term maintenance burden.

# What the author built: cxcap cxcap is an open-source, local, read-only CLI (cargo install cxcap) that aims to give a pre-change structural feedback loop. No index, config, daemon, model, or cloud account is required. Three primary commands cover common workflows:

  • cxcap audit. — show where complexity, hotspots, and cycles live
  • cxcap audit. --intent "add session expiry" — discover likely touchpoints for an intent
  • cxcap audit. --focus src/auth — show who depends on a target and how far its reach is

# Honest limitations reported by the author

# Practical takeaway for teams Continue to run tests and keep LLM review in your process. Add a local, deterministic pre-change audit as a complementary feedback loop: it asks a different question — what complexity will this change interact with? That question helps avoid accumulating structural costs that tests alone won't reveal.

More context around this story.

Your Free Probe Passed. You Still Measured the Wrong Box
Dev iconDevSep 13, 2026

Your Free Probe Passed. You Still Measured the Wrong Box

Your free probe passed on a scratch box. That green badge still measured the wrong machine. A green free endpoint is a cheap probe. It is not a production ship gate. Treat it like a probe, or you ship luck. I keep watching agents pass on free compute. Then the same patch dies on the real model. Does that gap sound fami

69 Tests. All Passing. Zero Bugs Caught.
Dev iconDevSep 17, 2026

69 Tests. All Passing. Zero Bugs Caught.

An AI model wrote 69 tests for a Python module. Every one passed. Together they caught zero of the eleven bugs I had deliberately planted in that module. A second setup, pointed at the specific bugs rather than at the module, used 17 attempts and caught ten. That contrast is the whole project. What I built A mutation-t

AI бЂ”бЂЉбЂєбЂёбЂ•бЂЉбЂ¬бЂЂбЂ­бЂЇ бЂЎбЂ™бЂјбЂ”бЂєбЂ†бЂЇбЂ¶бЂё бЂњбЂ±бЂ·бЂњбЂ¬бЂ”бЂЉбЂєбЂё
Medium iconMediumSep 5, 2026

AI бЂ”бЂЉбЂєбЂёбЂ•бЂЉбЂ¬бЂЂбЂ­бЂЇ бЂЎбЂ™бЂјбЂ”бЂєбЂ†бЂЇбЂ¶бЂё бЂњбЂ±бЂ·бЂњбЂ¬бЂ”бЂЉбЂєбЂё

AI (Artificial Intelligence) နည်းပညာက အá€á€¯á€¡á€á€»á€­á€”်မှာ နေရာá€á€­á€¯á€„်းမှာ ရှိနေပါပြီዠဒါပေမဲ့ “AI ကို ဘယ်ကနေ စလေ့လာရမလဲአအမြန်ဆုံး á€á€á€ºá€™á€¼á€±á€¬á€€á€ºá€¡á€±á€¬á€„်â

Loading more related stories...

Keep reading in the app

Open the app view to save this story, compare related coverage, and continue from the same source.

Open in app