Strata Runs a 125B AI Model on a Regular Gaming PC
Strata is a purpose-built inference engine that runs Qwen3.8-Flash-Next, a 125B mixture-of-experts model, on a single gaming GPU with 64GB of system RAM.
- Strata runs the 125B-parameter Qwen3.8-Flash-Next MoE on a single 12-24 GB NVIDIA GPU with 64 GB RAM.
- Users report around 70 tok/s on an RTX 3090 with 128K context using IQ2_XS quant.
- Serves OpenAI and Anthropic-compatible APIs on localhost 8080, drop-in for existing agents.
- Uses adaptive expert cache: GPU holds hot experts, CPU computes cold ones in parallel via AVX-512/AVX2.
- MTP speculative decoding accepts 2.4-3.2 tokens per pass while staying bit-identical to greedy decoding.
- Limits: single request at a time, greedy only, no KV reuse across turns, slow prefill at 128K+ contexts.
Strata serves a 125B MoE from a consumer PC
Strata is a new open-source inference server built specifically for Qwen3.8-Flash-Next, a 125-billion-parameter mixture-of-experts model that activates about 6 billion parameters for each token. The project runs quantized versions on a desktop with one NVIDIA GPU, 12 to 24 GB of VRAM, 64 GB of system memory, and an x86 CPU. It exposes OpenAI-compatible and Anthropic-compatible APIs on localhost.
The project’s published benchmarks show up to 94.6 tokens per second on an RTX 5070 at a 4K context and 65.1 tokens per second at 128K using its smallest quantization. An early user report on X cited about 70 tokens per second at 128K on an RTX 3090. Independent results across hardware configurations remain limited, and the project labels its 3090 and 40-series figures as estimates.
One model, three memory tiers
Qwen3.8-Flash-Next uses a mixture-of-experts architecture, which routes each token through a small subset of the model’s experts. All 125 billion parameters must remain available, but only about 6 billion participate in each token’s computation. Strata exploits that sparse activation pattern by treating GPU memory, system RAM, and SSD storage as one memory hierarchy.
- GPU: VRAM holds attention and DeltaNet mixer layers, routers, shared experts, the output head, the multi-token-prediction layer, the key-value cache, and an adaptive cache of frequently used experts.
- System RAM: All 24,576 experts remain pinned in memory. AVX-512 or AVX2 CPU kernels evaluate experts missing from the GPU cache while the GPU processes cached experts.
- SSD: A 28.8 GB n-gram table stays on disk. Strata reads a few rows per token through the operating system’s page cache.
During decoding, the model’s multi-token-prediction layer drafts as many as three tokens. One pass through all 48 layers verifies the draft, producing an average of 2.4 to 3.2 accepted output tokens per pass, according to the
This story is for Pro members
You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.