Batching by Length Instead of Looping Item by Item for SLM Optimization
Processing one short-text item per forward pass wastes memory bandwidth and time. Sort inputs by token length and form batches that pad only to the batch’s local maximum to cut redundant computation on padding tokens.




