Kdnuggets iconKdnuggetsSep 25, 2026 ~7 min source read

Batching by Length Instead of Looping Item by Item for SLM Optimization

Processing one short-text item per forward pass wastes memory bandwidth and time. Sort inputs by token length and form batches that pad only to the batch’s local maximum to cut redundant computation on padding tokens.

Batching by Length Instead of Looping Item by Item for SLM Optimization

Share this story

Send the public story page.

Useful takeaways from this story.

Reproducible setup: Qwen2.5-0.5B-Instruct, float16 via Hugging Face Transformers, M2 MacBook Air (24GB RAM, 16-core Neural Engine).

Simple implementation steps: tokenize, sort by length, create batches of similarly sized sequences, and request only the final logits when supported.

Batching amortizes weight reads across many sequences, but naive batching introduces another cost: padding. If you pad every batch to the global maximum sequence length in your dataset, most tokens in many batches are padding. Real text length distributions are long-tailed: a few long items and many short ones. Padding every batch to the dataset maximum causes a lot of unnecessary compute.

Sort tokenized examples by length before forming batches, so items grouped together have similar lengths and each batch only pads to its local maximum. This keeps the benefits of amortized weight reads while minimizing useless work on padding tokens.

  • Model: Qwen2.5-0.5B-Instruct.
  • Runtime: Hugging Face Transformers, float16, tested on an M2 MacBook Air with 24GB RAM and a 16-core Neural Engine.
  • Requirements: pip install torch transformers accelerate.
  • Implementation detail: set the tokenizer pad_token to eos if missing, use left padding so the final real token stays at index -1, and, where supported, request only the last logit to avoid allocating logits for every position.

The article simulates a realistic long-tailed ticket length distribution and runs each item as a single forward pass. Reported per-item timings in the baseline cluster around 0.23–0.24 s per ticket for increasing counts up to 600. Padding every item to the global maximum length would process 3.7× the necessary tokens, according to the example distribution used.

  • Tokenize every example and record its token length.
  • Sort the dataset by length (ascending or descending).
  • Form batches by taking contiguous runs of similarly sized items. Choose a batch size that matches your device memory budget.
  • Use left padding to keep the semantic content aligned at the right end of sequences when the model expects that.
  • If the model API supports returning only the final logits (or allows restricting logit positions), request that to reduce memory use for batched runs.

Grouping sequences of similar token length into batches reduces padding overhead and converts per-item weight streaming into a single amortized read, improving throughput on small models and low-memory hardware. The article demonstrates this with Qwen2.5-0.5B-Instruct and reports a 3.7× reduction in processed tokens compared with global-max padding in the example distribution.

More context around this story.

LLM Streaming with Embabel
Javacodegeeks iconJavacodegeeksSep 25, 2026

LLM Streaming with Embabel

Large Language Models can take several seconds, or sometimes much longer, to produce a complete response. In a traditional blocking application, the caller waits until the entire response has been generated before receiving anything. Streaming changes this interaction model. Instead of waiting for the complete LLM resp

Loading more related stories...

Keep reading in the app

Open the app view to save this story, compare related coverage, and continue from the same source.

Open in app