| |
Show HN: Run an 80B Qwen in 4.3 GB of RAM on a Mac, and a 35B on an iPhone
Swiftlet is a Swift + Metal runtime that enables running large Qwen AI models on Apple devices by streaming Mixture-of-Experts weights from storage on demand, allowing an 80B parameter model to run on a Mac with just 4.3 GB of RAM peak usage and a 35B model on an iPhone with approximately 2.5 GB of RAM. The trade-off is that only about 3B parameters are active per token, so the models perform like large models in conversation but have limited factual recall like smaller models, with decode speeds ranging from 1 token/second on iPhone to 4.5-11 tokens/second on Mac depending on the model size.
Read Full Article →
← More Tech news