Marktechpost iconMarktechpostSep 24, 2026 ~5 min source read

BottleCap AI Releases ThinkingCap-Qwen3.8-27B: 37.2% Fewer Thinking Tokens with a 0.86pp Accuracy Drop

ThinkingCap-Qwen3.8-27B is a fine-tune of Qwen3.8-27B focused on shorter reasoning traces. It reduces average reasoning tokens by 37.2% across 12 benchmarks while lowering macro accuracy by 0.86 percentage points and improving some long-context tasks.

BottleCap AI Releases ThinkingCap-Qwen3.8-27B: 37.2% Fewer Thinking Tokens at a 0.86pp Accuracy Cost

Share this story

Send the public story page.

Useful takeaways from this story.

With a solid foundation in statistical analysis, machine learning, and data engineering, Michal excels at transforming complex datasets into actionable insights.

Under a 16K-token cap per response, ThinkingCap scores higher than the base model.

BottleCap AI Releases ThinkingCap-Qwen3.8-27B: 37.2% Fewer Thinking Tokens at a 0.86pp Accuracy Cost.

# What ThinkingCap-Qwen3.8-27B Does

BottleCap AI fine-tuned Qwen3.8-27B to produce shorter reasoning traces without changing answer style or adding knowledge. The explicit goal: reduce unnecessary "thinking" tokens that the model emits during multi-step reasoning while keeping instruction following, safety behavior, and core reasoning ability largely intact.

# Main results at a glance

# Benchmark-level trade-offs

# Effort dial interaction

Qwen3.8-27B exposes a reasoning-effort setting. BottleCap reports that ThinkingCap's compression stacks with that dial. Compared to the base model at xhigh, ThinkingCap cuts thinking tokens more aggressively across medium and low effort settings while keeping accuracy shifts similar. With thinking disabled altogether, ThinkingCap trails the base by 5.7pp. BottleCap recommends xhigh for the best accuracy-to-token balance and plans future work on per-mode tuning.

# Evaluation setup and reproducibility

# Deployment and builds

The bf16 checkpoint contains 28B parameters and accepts image and text inputs. BottleCap provides five quantized builds:

  • FP8: 31 GB, supported on vLLM for Hopper and Blackwell GPUs.
  • NVFP4 weight-only: 21 GB, vLLM on Hopper (Marlin kernel) and Blackwell.
  • NVFP4 W4A4 (AWQ): 23 GB, Blackwell only.
  • GGUF: 16–55 GB, for llama.cpp, LM Studio and Ollama.
  • MLX 4-bit DWQ: 21 GB, targeted at Apple Silicon Macs with 32 GB.

Serving follows the base model recipe with --reasoning-parser qwen3 and the qwen3_xml tool-call parser on vLLM. Thinking outputs appear in a separate reasoning field.

# License and access

Weights are gated. The default license is PolyForm Small Business 1.0.0 plus a BottleCap personal-use grant. Upstream Qwen materials remain under Apache-2.0. Commercial use beyond the small-business license requires a BottleCap agreement. No inference-provider hosting was listed on Hugging Face at the time of reporting.

# Practical implications

If you need to reduce token consumption for cost or latency reasons while accepting small accuracy losses on average, ThinkingCap offers a straightforward substitute for Qwen3.8-27B on vLLM and SGLang. Task-specific evaluation is necessary: some benchmarks maintain accuracy with large token savings, while a few (notably AIME 2026) show meaningful accuracy drops.

More context around this story.

What is agentic AI in retail?
Legaltechdaily iconLegaltechdailySep 8, 2026

What is agentic AI in retail?

Your customers expect consistent, instant, intelligent service across every channel, even when they bounce between SMS, email, and web chat from one interaction to the next. But your team can’t be everywhere at once—no one can. That’s exactly where agentic AI in retail comes in. Agentic AI gives customers a seamless ex

Loading more related stories...

Keep reading in the app

Open the app view to save this story, compare related coverage, and continue from the same source.

Open in app