Dzone iconDzoneSep 28, 2026 ~8 min source read

Stop Paying a Model to Re‑decide Procedures Your Team Already Set

When teams distribute reusable developer skills, those skills can cause repeated model round trips by reconstructing fixed procedures as prose. Measure the cost with the vendor telemetry already available, add a small privacy-preserving hook for command context, and restructure work so the model isn’t re‑deciding what the system has already decided.

Stop Paying a Model to Make Decisions You Already Made

Share this story

Send the public story page.

Useful takeaways from this story.

Use the vendor’s OpenTelemetry export to attribute inference round trips to specific skills via the skill.name field, then measure cost in API requests and cost_usd estimates.

Enable OTEL_LOG_TOOL_DETAILS=1 and add a small local hook that emits a privacy-preserving shape of commands joined by tool_use_id to correlate tool results with telemetry.

# Why this matters Teams that package developer skills (for example as a Claude Code skill, a Cursor plugin, or a shared rules file) create a catalog of how work should be done: service structure, commit gates, which internal library to use. Those artifacts are useful, but if a skill spells out a fixed procedure in prose, the agent will reconstruct and deliberate over that procedure every invocation. That procedural amplification shows up as extra model round trips and higher inference costs even when the context size looks small.

# What you can measure today Claude Code's OpenTelemetry export already includes event-level data that matters: api_request events carry skill.name, cost_usd (an estimate), input_tokens, output_tokens, duration_ms, and an event.sequence. Tool_result events include tool_use_id, tool_name, success, and duration_ms. Crucially, api_request events attribute inference round trips to the active skill via skill.name, so you can count round trips per skill and convert to cost estimates without building a separate pipeline.

# Capture the one missing piece: command shape

  • Transform commands with an allowlist of binaries and a shape map that replaces argument values with placeholders. Example: git commit -m "fix auth for acme" becomes git commit -m.
  • Track flags that consume the next token (for example --token) so you don't accidentally leak a bare secret. The safety of the pipeline depends on getting the allowlist right.
  • Emit only the join key and the shaped command. Do not duplicate fields that OTel already reports (success, duration_ms, skill.name, token counts).

# Hook performance and placement

  • Subprocess hook: p50 ≈ 24 ms, p99 ≈ 32 ms, mean ≈ 25 ms
  • In‑process hook: p50 ≈ 0.08 ms, p99 ≈ 0.20 ms, mean ≈ 0.10 ms

On macOS arm64 the in‑process approach is orders of magnitude cheaper. If a hook must be slower, prefer async execution so the agent doesn't wait for it.

# How to decide whether a skill should be procedural or declarative

# Actionable checklist

  • Enable vendor OTel export and collect api_request and tool_result events.
  • Set OTEL_LOG_TOOL_DETAILS=1 to get tool-level command fields.
  • Add a small local hook that emits tool_use_id plus a privacy‑preserving shaped command using an explicit allowlist.
  • Identify skills with high round‑trip counts via skill.name and consider replacing procedural prose with deterministic checks or minimal code.

# Bottom line Measure inference cost per skill using the telemetry you already have, capture only the minimal safe command context locally, and stop paying for the model to re‑decide steps the system can decide deterministically.

More context around this story.

Decision Dependency

Decision Dependency

When something doesn’t make sense financially, business owners naturally turn to the accounting records for answers.Sometimes the accounting records provide the answer. Sometimes they simply tell us where the problem became visible.

Prompt Caching Doesn't Save Money on Turn One
Dzone iconDzoneSep 25, 2026

Prompt Caching Doesn't Save Money on Turn One

I went looking for a clean way to show what prompt caching actually saves an agent, and the first thing I found was a fact that's easy to miss if you only read the "up to 90% savings" headline. Caching a fresh conversation's first turn costs more than not caching it. There's no cache to read from yet, so you pay the in

How to Avoid APM Bill Surprises
Fastruby iconFastrubySep 1, 2026

How to Avoid APM Bill Surprises

Originally appeared on The Rails Tech Debt Blog . You approved an APM tool at a modest monthly rate. A year later, the renewal invoice bears little resemblance to what you signed, and nobody remembers deciding to spend that much more. APM is worth having when it’s used well: it resolves incidents faster by showing you

Keep reading in the app

Open the app view to save this story, compare related coverage, and continue from the same source.

Open in app