Tooling
Qwen 27B reaches 218 tokens/sec on dual RTX 3090s with vLLM
A developer achieved 218 tokens per second decode throughput running Qwen 27B on two consumer GPUs using vLLM, DFlash2 speculative decoding, and INT4 quantization.
1 min read
Sourcer/localllama
A developer running Qwen 27B on dual RTX 3090 GPUs achieved 218 tokens per second decode throughput using vLLM v0.26.1rc1, DFlash2 speculative decoding, and INT4 quantization. The result, measured with the Club-3090 canonical benchmark suite, demo...
Sign in to read the full analysis
Free account. Full analysis on LLM unit economics, plus the weekly Cost-of-Inference column.
Try it on your own context
You just read the writeup. Now run the thing. Paste a doc or some verbose tool output and watch it shrink — free, no signup.
2,912/12,000 chars
Compressed
Compressed text will appear here…
Method & sources
- Source type
- Primary publication (lab/vendor blog) — our analysis + implication
- Source link
- r/localllama
- Published
- UTC
- Byline
- By the gotcontext.ai team (editorial standards)
- Correction?
- corrections@gotcontext.ai