Tooling
Developer releases 250M quantized LLM in 60 MB footprint
A developer trained a 250M parameter language model on 30B tokens and compressed it to under 2 bits, achieving a 60 MB deployment size that runs on CPU without GPU requirements.
1 min read
Sourcer/localllama
A developer has released SHADOW-250M, a 250M parameter language model trained from scratch on 30B tokens of fineweb data and quantized to under 2 bits, resulting in a 60 MB deployment footprint that runs on standard laptop CPUs at approximately 400 tokens per second. The model requires only 80 MB of...
Sign in to read the full analysis
Free account. Full analysis on LLM unit economics, plus the weekly Cost-of-Inference column.
Try it on your own context
You just read the writeup. Now run the thing. Paste a doc or some verbose tool output and watch it shrink — free, no signup.
2,912/12,000 chars
Compressed
Compressed text will appear here…
Method & sources
- Source type
- Primary publication (lab/vendor blog) — our analysis + implication
- Source link
- r/localllama
- Published
- UTC
- Byline
- By the gotcontext.ai team (editorial standards)
- Correction?
- corrections@gotcontext.ai