Cactus Compute releases Needle 3, an 8–29MB automation model that matches DeepSeek V4 Flash on tool‑call and structured‑JSON tasks. The architecture uses a Monarch Hadamard MLP and intelligent layer‑laddering to deliver up to 4 k tokens/sec on a Raspberry Pi 5 while supporting eight languages and easy fine‑tuning for narrow domains.
Background
There is growing demand for lightweight, on‑device AI models that can perform specialized automation tasks without relying on cloud APIs. Needle 3 contributes to this trend by demonstrating that careful architectural compression and intelligent subnetwork selection can achieve performance close to much larger commercial models.
- Source
- Hacker News (RSS)
- Published
- Sep 18, 2026 at 08:11 AM
- Score
- 9.0 / 10