A developer built SlotStream, a Mac-native tool using MLX and Swift that streams a 125B-parameter Qwen3.8-Flash-Next model (4-bit, ~104GB) onto a 48GB Mac at ~12 tok/s via expert-offloading and SSD-streaming. The auto-mode balances memory and speed, and the roadmap includes MTP-based speculative decoding for further acceleration.
Background
Running large language models on consumer hardware with limited RAM remains a major challenge; techniques like offloading and SSD streaming are gaining attention as practical solutions.
- Source
- Hacker News (RSS)
- Published
- Sep 2, 2026 at 12:42 AM
- Score
- 7.0 / 10