Project Maya Runs Z.ai's 321B GLM-5.3-Flash on Two Consumer GPUs and an SSD
Project Maya runs the 321B GLM-5.3-Flash mixture-of-experts model on one or two consumer NVIDIA GPUs by tiering experts across VRAM, RAM, and SSD.
- Project Maya runs the 321B-parameter GLM-5.3-Flash MoE on 1-2 NVIDIA GPUs by tiering experts across VRAM, RAM, and NVMe SSD.
- The Maya-S quant fits the whole model in 96.5 GB using error-feedback rounding from Z.ai's FP8 weights.
- Reported throughput: up to 40 tokens/s generation and 440 tokens/s prompt on 2x Tesla V100 32GB with 30 GB RAM.
- Zero-shot accuracy is 97.9% of the full FP8 model; token agreement is 83.3%, weakest on tool calls.
- Minimums: 32 GB RAM, ~100 GB NVMe, Linux + CUDA 12.x; Windows and WSL2 are not yet supported.
- Serves OpenAI and Anthropic compatible APIs, includes vision, MTP speculative decoding, and a one-command installer.
Project Maya runs a 321B MoE across two V100s, RAM, and NVMe
Project Maya combines a custom inference engine with a heavily quantized version of Z.ai’s GLM-5.3-Flash. The project’s reference setup runs the 321-billion-parameter model on two 32 GB Tesla V100 GPUs, system memory, and a fast NVMe SSD.
The stack has two parts. The quantized model, called Maya-S, stores the weights in a 96.5 GB GGUF file. Project Maya’s engine distributes those weights across available VRAM, RAM, and storage during inference.
GLM-5.3-Flash uses a mixture-of-experts architecture. It contains 321 billion parameters in total, but activates about 18 billion for each token. That sparse activation pattern lets Maya keep frequently selected experts near the GPUs while leaving less-used weights on slower storage.
Hot experts on GPU, cold experts on SSD
Maya ranks experts by use and assigns them to three tiers. Frequently selected experts stay in VRAM, the next group occupies system memory, and the remaining weights remain on NVMe until requested. The cache adapts as prompts exercise different parts of the model.
The runtime profiles the machine during startup. It fills available VRAM, divides layers across two GPUs when present, sizes the RAM tier from free memory, and compares CPU performance with PCIe transfer speed. That benchmark determines whether RAM-resident experts should run on the CPU or move to a GPU for computation.
This design trades storage traffic and cold-start latency for lower memory requirements. Initial responses run more slowly while the expert caches warm. Performance then depends on how consistently later prompts reuse the cached experts.
The practical hardware floor
The maintainers report up to 40 generated tokens per second and 440 prompt tokens per second on the reference machine. The generation result uses the model’s NextN multi-token prediction layer for speculative decoding, which drafts several tokens and verifies them together.
This story is for Pro members
You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.