E-Ink News Daily

Back to list

Bonsai 2 27B: Near-Lossless Compression in a 9x Smaller Footprint

PrismML released Ternary Bonsai 2 27B, a near-losslessly compressed version of Qwen3.8 27B using ternary weights ({−1, 0, +1}) with FP16 group-wise scaling, achieving 1.76 effective bits per weight and a 5.9GB footprint. The model retains 98.2% of its full-precision counterpart's benchmark performance while being over 9x smaller, supporting a 262K-token context window, multimodal input, and agentic capabilities under Apache 2.0.

Background

Model compression and quantization are critical for deploying large AI models on edge devices and local hardware. Ternary quantization reduces weights to three values (−1, 0, +1), dramatically shrinking model size while attempting to preserve performance.

Source
Lobsters
Published
Sep 18, 2026 at 04:28 PM
Score
7.0 / 10