# What ThinkingCap-Qwen3.8-27B Does
BottleCap AI fine-tuned Qwen3.8-27B to produce shorter reasoning traces without changing answer style or adding knowledge. The explicit goal: reduce unnecessary "thinking" tokens that the model emits during multi-step reasoning while keeping instruction following, safety behavior, and core reasoning ability largely intact.
# Main results at a glance
# Benchmark-level trade-offs
# Effort dial interaction
Qwen3.8-27B exposes a reasoning-effort setting. BottleCap reports that ThinkingCap's compression stacks with that dial. Compared to the base model at xhigh, ThinkingCap cuts thinking tokens more aggressively across medium and low effort settings while keeping accuracy shifts similar. With thinking disabled altogether, ThinkingCap trails the base by 5.7pp. BottleCap recommends xhigh for the best accuracy-to-token balance and plans future work on per-mode tuning.
# Evaluation setup and reproducibility
# Deployment and builds
The bf16 checkpoint contains 28B parameters and accepts image and text inputs. BottleCap provides five quantized builds:
- FP8: 31 GB, supported on vLLM for Hopper and Blackwell GPUs.
- NVFP4 weight-only: 21 GB, vLLM on Hopper (Marlin kernel) and Blackwell.
- NVFP4 W4A4 (AWQ): 23 GB, Blackwell only.
- GGUF: 16–55 GB, for llama.cpp, LM Studio and Ollama.
- MLX 4-bit DWQ: 21 GB, targeted at Apple Silicon Macs with 32 GB.
Serving follows the base model recipe with --reasoning-parser qwen3 and the qwen3_xml tool-call parser on vLLM. Thinking outputs appear in a separate reasoning field.
# License and access
Weights are gated. The default license is PolyForm Small Business 1.0.0 plus a BottleCap personal-use grant. Upstream Qwen materials remain under Apache-2.0. Commercial use beyond the small-business license requires a BottleCap agreement. No inference-provider hosting was listed on Hugging Face at the time of reporting.
# Practical implications
If you need to reduce token consumption for cost or latency reasons while accepting small accuracy losses on average, ThinkingCap offers a straightforward substitute for Qwen3.8-27B on vLLM and SGLang. Task-specific evaluation is necessary: some benchmarks maintain accuracy with large token savings, while a few (notably AIME 2026) show meaningful accuracy drops.