E-Ink News Daily

← Back to list

Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s

Strata is an open-source inference engine that enables running the massive 125B-parameter Qwen3.8-Flash-Next model on consumer NVIDIA or AMD GPUs (12GB+ VRAM) via a one-click installer for Windows/Linux. It provides a local OpenAI/Anthropic-compatible API and supports image input, achieving up to 100 tokens/s on an RTX 4090.

Background

The demand for accessible local LLM inference has grown rapidly, with quantization techniques like GGML enabling large models to run on consumer hardware. This project follows the trend of democratizing AI by making billion-parameter models available outside data center environments.

Source
hackernews
Published
Oct 4, 2026 at 08:51 PM
Score
7.0 / 10