Models
Ling 3.0 Tiny reaches 36 tokens per second on 4GB VRAM
Ling 3.0 Tiny, an 8-billion parameter model with only 1.3 billion active parameters, delivers 36 tokens per second on consumer hardware with 4GB of VRAM.
1 min read
Sourcer/localllama
Ling 3.0 Tiny, an 8-billion parameter model with 1.3 billion active parameters, achieves 36 tokens per second on a system with 4GB of VRAM. The model combines sparse activation with aggressive quantization to deliver inference speeds while maintaining competitive reasoning performance.
According to...
Sign in to read the full analysis
Free account. Full analysis on LLM unit economics, plus the weekly Cost-of-Inference column.
Try it on your own context
You just read the writeup. Now run the thing. Paste a doc or some verbose tool output and watch it shrink — free, no signup.
2,912/12,000 chars
Compressed
Compressed text will appear here…
Method & sources
- Source type
- Primary publication (lab/vendor blog) — our analysis + implication
- Source link
- r/localllama
- Published
- UTC
- Byline
- By the gotcontext.ai team (editorial standards)
- Correction?
- corrections@gotcontext.ai