PrismML, a startup founded by Caltech researchers, is developing highly compressed reasoning LLMs that can run on PCs and smartphones. Their latest model, Bonsai 2 27B, compresses Alibaba's Qwen3.8 27B down to just 5.9 GB — a 9x to 10x memory reduction. The company, advised by Databricks co-founder Ion Stoica and backed by Khosla Ventures, is rumored to be in talks with Apple.
Background
The push for efficient, on-device AI models has accelerated as companies seek to reduce cloud dependency and improve privacy. LLM compression and quantization have become active research areas across the industry.
- Source
- TechCrunch
- Published
- Sep 18, 2026 at 06:34 AM
- Score
- 7.0 / 10