A project called Deltafin demonstrates running the massive Kimi K3 2.8T parameter model on a MacBook Pro at 1 token/s by streaming weights from four SSDs. This showcases impressive optimization for local inference of extremely large language models on consumer hardware.
Background
Large language models with trillions of parameters typically require massive GPU clusters to run. Streaming model weights from SSDs to CPU/RAM is an emerging technique to run huge models on consumer hardware.
- Source
- Hacker News (RSS)
- Published
- Sep 9, 2026 at 04:07 AM
- Score
- 7.0 / 10