Medium iconMediumAug 24, 2026

一張 4090 跑 Qwen3.8-27B:vLLM vs llama.cpp 實測,和一個沒人在講的免費 1.5 倍 context

Qwen3.8–27B 現在是很多人本機跑 coding model 的首選,而大部分人第一個拿起來的工具是 llama.cpp — 簡單、GGUF 原生、跑得動。我們把兩套 stack 放在同一張 RTX 4090 上跑,原本預期 vLLM 會在 agent… Continue reading on Medium »

Share this story

Send the public story page.

Useful takeaways from this story.

Qwen3.8–27B 現在是很多人本機跑 coding model 的首選,而大部分人第一個拿起來的工具是 llama.cpp — 簡單、GGUF 原生、跑得動。我們把兩套 stack 放在同一張 RTX 4090 上跑,原本預期 vLLM 會在 agent… Continue reading on Medium »

Building the complete brief

The page is ready to read now. The fuller skim-friendly version will appear here automatically.

The useful part

Qwen3.8–27B 現在是很多人本機跑 coding model 的首選,而大部分人第一個拿起來的工具是 llama.cpp — 簡單、GGUF 原生、跑得動。我們把兩套 stack 放在同一張 RTX 4090 上跑,原本預期 vLLM 會在 agent… Continue reading on Medium »

Keep reading in the app

Open the app view to save this story, compare related coverage, and continue from the same source.

Open in app