Strata Runs a 125B-Parameter Qwen Model on a 12GB Gaming GPU
Strata is an MIT-licensed inference engine that runs Qwen3.8-Flash-Next, a 125-billion-parameter Mixture-of-Experts model with about 6 billion active per token, on consumer GPUs with 12GB of VRAM. It caches the most-used experts in VRAM, keeps the rest in system RAM, and reached 94 tokens per second on an RTX 5070 in the project's tests.
On this page
Expert hot caching, explained
An open-source project called Strata is trending for a claim that would have sounded implausible a year ago: a 125-billion-parameter model running on the kind of graphics card gamers already own. The engine is MIT-licensed, installs with one click on Windows or Linux, keeps everything on your machine, and serves an OpenAI-compatible endpoint on localhost alongside Anthropic and Responses APIs plus an MCP server for agents.
The trick is that the model, Alibaba's Qwen3.8-Flash-Next, is an ultra-sparse Mixture-of-Experts: 125 billion total parameters organized as 24,576 specialist experts, of which each token only activates about ten, roughly 6 billion parameters of work. Strata exploits that sparsity with expert hot caching: the GPU's 12GB of VRAM holds the few thousand most frequently used experts, system RAM holds the complete set while the CPU computes the overflow, and the SSD stores a lookup table for the cold tail. A speculative decoding pass, where a small model drafts and the big model verifies in parallel, cuts latency another 1.6 to 1.8 times.
The numbers on real hardware
On an RTX 5070 with 12GB, the project's own benchmarks show Q2_0 quantization writing about 94 tokens per second and IQ3_S about 53, with prompt reading at 1,600 to 2,650 tokens per second; on AMD, an RX 9070 XT writes around 60 tokens per second. A discussion thread on Hacker News, where a front-page post about running the model on an RTX 4090 drew more than 300 points, suggests real users are getting similar results on 4090-class cards at around 100 tokens per second.
The requirements are honest about the trade. The GPU needs 12GB or more, but the machine also wants 32GB of system RAM as a floor, with 64GB recommended because the engine loads 35 to 55GB of experts into memory, plus roughly 80GB of SSD space for the model. First startup can freeze the PC for a few minutes, requests are handled serially by default, and weak spots exist: the Coder variant struggles with CJK text, and image input on AMD cards is Linux-only.
What it means for local AI
The significant shift is which barrier fell. Local model enthusiasts have long known that ultra-sparse MoEs do most of their work in a small slice of the network; Strata's contribution is engineering that insight into an installer that a non-expert can run on yesterday's gaming PC. The model it serves is itself open-weight and available through Ollama, so the full stack, weights, engine, and API, sits under user control.
For privacy-conscious users the practical upshot is simple: a frontier-class open model that once demanded a data center GPU now runs entirely on hardware that fits under a desk, with no network calls at all. Expect the next round of this race to be caching strategy versus context length, because the same trick that makes 125 billion parameters fit into 12GB of VRAM will be tested by everything else people want their local agents to remember.