Count Tokens and Estimate Your LLM API Costs Before You Ship
Your bill is the product of two measurable token counts and two rates. Most surprises come from how text tokenizes and from repeatedly sending conversation history.

Your bill is the product of two measurable token counts and two rates. Most surprises come from how text tokenizes and from repeatedly sending conversation history.

Every request includes ~7–9 tokens of envelope overhead beyond visible message content.
Output tokens cost significantly more than input tokens, so reply-heavy systems cost more than input-heavy ones.
# What this guide explains
# What a token actually is A token is a chunk of text the model treats as one unit. Common words are often one token. Rare words, punctuation, spaces, line breaks, JSON, code, UUIDs and other identifiers can split into many tokens. Tokenization depends on the content, not simply length. That is why rules of thumb based on characters or words drift widely across real inputs.
# The rule of four characters per token is unreliable
# How to count reliably
# Watch out for envelope tokens Counting only visible message text undercounts. Real API requests include structural tokens for roles and message framing. Measured requests show roughly seven to nine tokens of envelope per request. That overhead is small for very large prompts but dominates the cost of many tiny requests.
# Output costs and cost shape Providers charge more for generated output tokens than for input tokens. Across major providers measured, output costs were two to six times higher than input. That means a system that produces long replies will have a different cost profile than one that performs short summaries or tagging. Measure both input and expected output token counts to model costs accurately.
# Conversation history and the common surprise Teams commonly discover large bills when a system keeps resending conversation history. Ten conversation turns can represent about 5.1× the text that actually exists in a single turn, because history accumulates every request. This growth is the most frequent source of unexpected spend.
# Prompt caching and discounts Prompt caching is an available discount: cached reads are billed at roughly a tenth of normal input costs, according to measured behavior. Prompt caching is the largest effective discount available, but many systems do not use or claim it. Where caching makes sense, it sharply reduces per-request input cost.
# Practical next steps 1) Tokenize representative prompts, system messages, and likely user inputs with the same tokenizer the model uses. 2) Measure input and output token counts separately. 3) Include envelope tokens per request in your model. 4) Model conversation growth over expected turns to estimate accumulation. 5) Implement prompt caching where appropriate to cut input cost.
# Final point The usage field returned by API responses provides exact billed counts. After one real call, you can replace estimates with measured values and budget accurately.
Before diving into solutions, it helps to understand the scale of the problem. Take a common production pattern: a customer support bot that processes 10,000 messages per day, each with a 2,000-token system prompt and a 200-token user message.

AI token spend for business grew 572% year over year from June 2025 to June 2026, according to proprietary data from Ramp, which processes AI vendor payments on behalf of thousands of businesses. Finance teams are catching up to what engineering already knows: AI isn’t a line item you set once and forget. It’s

Splunk’s open-source Token Meter gives developers real-time visibility into AI coding agent activity, token consumption and estimated costs across tools including Claude Code, Codex and Cursor.

The Amazon Bedrock pricing page publishes a training price for the Meta model you almost certainly are not fine-tuning. Here is the customization pricing on the page, as of 8 September 2026: Model Published training price Llama 2 Pretrained 13B $1.49 per 1M tokens Llama 2 Pretrained 70B $7.99 per 1M tokens Cohere Comma
После недели работы с кодинг-агентом возникает вопрос: почему счёт растёт быстрее, чем объём написанного кода. Разбираю по полям usage, куда именно уходят токены. Оказывается, типов не два, а пять: обычный ввод, запись в кэш, чтение из кэша, вывод и reasoning. Последний виден не всегда, а платите вы за него по цене выв

Originally appeared on All about coding . LLMs work with the numbers behind your text rather than the text itself. Two Rails class definitions one word apart encode to arrays of 5 token IDs that differ in a single position, so subtracting one from the other returns the class name.
Loading more related stories...
Open the app view to save this story, compare related coverage, and continue from the same source.