There’s a new project on Hacker News today that puts 125 billion parameters on a gaming GPU. The benchmark number is 94 tokens per second on an RTX 5070 with 12GB of VRAM. That headline is technically accurate, and also somewhat misleading, and the gap between those two things is the interesting part.
The model in question is Qwen3.8-Flash-Next, Alibaba’s August 2026 release — a 125B MoE with 512 experts, of which only 10 are routed per token (plus 1 shared), giving roughly 6B active parameters per forward pass. There’s also a 51B n-gram embedding layer and a 4B multi-token prediction head that factor into total parameter count but not into the core inference compute. The practical implication: each token visits a tiny fraction of the total weights. Most of the model is idle at any given moment.
Strata exploits this directly. It organizes the experts into a tiered memory hierarchy: frequently-accessed experts live in GPU VRAM, all experts are pinned to system RAM, a parallel CPU thread handles any overflow, and lookup tables are stored on SSD. Because the active set per token is small and relatively stable — popular experts stay hot — VRAM thrash is limited. You need 80GB of disk and 32GB of RAM, but only 12GB of VRAM. The full model weight is distributed across storage tiers that differ by three orders of magnitude in bandwidth, and it mostly works because MoE sparsity keeps the working set narrow.
On top of that, Strata pairs the main model with a smaller helper for speculative decoding: the helper proposes a sequence of next tokens, the primary model validates all predictions in parallel, and accepted tokens arrive without additional latency. The project reports a 1.6–1.8× throughput gain from this stage alone.
The combination brings a model whose parameter count matches GPT-3-class density into the practical range of hardware a developer might already own. Quality-wise, Qwen3.8-Flash-Next is positioned as a preview of Qwen4’s architecture — Alibaba’s framing is “quality of a 27B dense model, compute of a 6B” — which puts it somewhere in the range of serious general-purpose capability.
None of this is entirely new. Expert offloading to RAM and speculative decoding are both established techniques. What’s notable here is the execution: a clean open-source implementation, MIT license, detailed hardware requirements documented upfront, and a project page that doesn’t oversell what’s happening. The HN title says “RTX 4090” in the headline but the README benchmarks show an RTX 5070 — a minor mismatch, but the kind of thing worth checking before citing. The performance is real; the exact hardware configuration matters.
The deeper point is about architecture selection. MoE was adopted at scale primarily because it allows training larger models without proportional compute cost — you get the parameters cheaply during training by only activating a fraction per sample. The inference story is different and, for local deployment, accidentally better: the same sparsity that makes training efficient also means most weights are cold during inference. That cold majority can live in slow storage. You can run a “125B” model locally not despite the MoE architecture but because of it.
What’s still not great: latency for short exchanges is higher than a dense 7B equivalent, the first-token time includes RAM-to-VRAM transfer overhead, and the system is sensitive to RAM bandwidth. It’s not a replacement for a fast cloud API for interactive use. But for offline workloads, document processing, and contexts where privacy or cost matter more than latency, the performance profile is now genuinely competitive.
Strata joins a cluster of projects — antirez’s DwarfStar, Magnitude’s device kernel engine — that are collectively making it harder to justify “this model requires cloud” as a default assumption.
