When agents pass tests but make future changes harder
Frontier coding agents can make locally correct edits that increase repository coupling and complexity. A local, deterministic pre-change complexity audit can reveal risks tests don’t measure.

Frontier coding agents can make locally correct edits that increase repository coupling and complexity. A local, deterministic pre-change complexity audit can reveal risks tests don’t measure.

Repeated agent-driven edits can accumulate coupling, duplication, and hotspots that raise future maintenance and agent-workload costs.
A read-only, local complexity map (cxcap) can surface hotspots, transitive exposure, and likely touchpoints before implementing a change.
# The blind spot: tests verify correctness, not structural impact Coding agents are effective at the immediate task in front of them: inspect local context, implement, test, fix, ship. That loop can show green tests and reasonable diffs, yet the repository quietly accumulates coupling, duplicated behavior, larger modules, and transitive exposure. Over many iterations this accumulation makes future changes harder.
# A concrete case: rounding was never just rounding The author built a synthetic repo (AcmeSaaS: auth, billing, notifications, API, workers — 83 files, Python and TypeScript) and gave an agent the task "change currency rounding." A single-line conceptual change turned out to touch many files. An audit run reported three likely touchpoints and a 15-file reasoning surface, and then revealed a surprising hotspot: core/money.py.
# External evidence that this pattern exists at scale The problem the author observed aligns with broader measurements. GitClear reported increases in block duplication (+81%), within-commit copy/paste (+41%), and error-masking constructs (+47%) across a dataset of 623 million code changes. DORA data shows AI adoption speeds creation, which reallocates time to auditing and verification and can raise delivery instability alongside increased output velocity. Those findings point to faster generation increasing the need for structural feedback.
# Why complexity matters economically Technical debt now has an operational cost beyond developer time. More complex repositories force agents to ingest more files, call more tools, and reason across more relationships. Agent platforms meter usage by tokens, model, context, and workload, so unnecessary structural complexity can raise recurring AI-workload overhead as well as long-term maintenance burden.
# What the author built: cxcap cxcap is an open-source, local, read-only CLI (cargo install cxcap) that aims to give a pre-change structural feedback loop. No index, config, daemon, model, or cloud account is required. Three primary commands cover common workflows:
# Honest limitations reported by the author
# Practical takeaway for teams Continue to run tests and keep LLM review in your process. Add a local, deterministic pre-change audit as a complementary feedback loop: it asks a different question — what complexity will this change interact with? That question helps avoid accumulating structural costs that tests alone won't reveal.
A few months ago, I saw something that made me rethink what coding assistants are actually capable of. A teammate was dealing with a frustrating race condition hidden deep inside a legacy service. It wasn't an obvious bug, and it had already taken quite a bit of time to investigate. Instead of digging through the code

Your free probe passed on a scratch box. That green badge still measured the wrong machine. A green free endpoint is a cheap probe. It is not a production ship gate. Treat it like a probe, or you ship luck. I keep watching agents pass on free compute. Then the same patch dies on the real model. Does that gap sound fami

An AI model wrote 69 tests for a Python module. Every one passed. Together they caught zero of the eleven bugs I had deliberately planted in that module. A second setup, pointed at the specific bugs rather than at the module, used 17 attempts and caught ten. That contrast is the whole project. What I built A mutation-t

A personal-assistant agent runtime protected by 4,286 unit tests and 827 declarative governance checks suffered 22 silent failures over… Continue reading on Medium В»
How to verify your app aligns with your intent without ever reading a line of generated code. The post Coding Agents Keep Shipping Silent Failures — Here Is How to Catch Them appeared first on Towards Data Science .

AI (Artificial Intelligence) နည်းပညာက အá€á€¯á€¡á€á€»á€á€”်မှာ နေရာá€á€á€¯á€„်းမှာ ရှá€á€”ေပါပြီዠဒါပေမဲ့ “AI ကá€á€¯ ဘယ်ကနေ စလေ့လာရမလဲአအမြန်ဆုံး á€á€á€ºá€™á€¼á€±á€¬á€€á€ºá€¡á€±á€¬á€„်â
Loading more related stories...
Open the app view to save this story, compare related coverage, and continue from the same source.