Swiftlet runs 80B Qwen model on 4.3 GB Mac RAM
Swiftlet, a new inference framework, enables running Alibaba's 80 billion parameter Qwen model on macOS with just 4.3 GB of RAM and a 35B variant on iPhone, using aggressive quantization and memory optimization.
Swiftlet, an open-source inference framework, demonstrates that large language models can run on consumer devices with extreme resource constraints. The project enables an 80 billion parameter Qwen model to execute on macOS using only 4.3 GB of RAM, while a 35 billion parameter variant runs on iPhon...
Sign in to read the full analysis
Free account. Full analysis on LLM unit economics, plus the weekly Cost-of-Inference column.
Try it on your own context
You just read the writeup. Now run the thing. Paste a doc or some verbose tool output and watch it shrink — free, no signup.
- Source type
- Primary publication (lab/vendor blog) — our analysis + implication
- Source link
- Hacker News · Front Page
- Published
- UTC
- Byline
- By the gotcontext.ai team (editorial standards)
- Correction?
- corrections@gotcontext.ai