China Telecom's Xing4.0 Runs a 29B Coding Agent on 19GB Locally

A Chinese Telecom lab released a 29B mixture-of-experts model with only 4B active parameters, 256K context, and strong agentic coding benchmarks.

·
·
·
China Telecom's Xing4.0 Runs a 29B Coding Agent on 19GB LocallyPRO
  • China Telecom released Xing4.0-29B-A4B, a 29B MoE with only 4B active parameters.
  • Native 256K context extensible to 512K using MLA attention and 64 routed experts.
  • Scores 75.0 on SWE-bench Verified and 57.5 on Terminal-Bench 2.1.
  • First model at this scale trained end-to-end on Huawei Ascend 910C NPUs with MindSpore.
  • GGUF quantizations from 9.9 GB (IQ2_M) to 62.5 GB (F16); Apache 2.0 licensed.
  • See the training report and GitHub repo for details.

Xing4.0 brings a 29B mixture-of-experts model to GGUF

A GGUF conversion of Xing4.0-29B-A4B has been published on Hugging Face. Developed by China Telecom’s AI team, whose earlier work includes TeleChat, the model contains 29 billion parameters while activating about 4 billion for each token. The release targets local coding agents, repository analysis, and long-context research workloads.

29B stored, 4B routed

Xing4.0 uses a sparse mixture-of-experts architecture with 40 layers and a hidden size of 3,584. Each token is routed through four of 64 specialized experts plus one shared expert. This routing reduces computation per token, while all 29 billion parameters must remain available in system memory, GPU memory, or a combination of both.

Multi-head Latent Attention, commonly shortened to MLA, compresses the key-value cache used to track earlier tokens. The model card lists a native context window of 256,000 tokens and extension up to 512,000. MLA makes those windows more practical, although memory use still rises with context length, batch size, and cache precision.

Quantization sets the hardware floor

The GGUF repository provides several quantizations that trade file size against numerical precision:

Build Approximate size Use case
IQ2_M 9.9 GB Maximum compression for limited memory
Q4_K_M 19 GB Repository-recommended balance
Q6_K 25.7 GB Lower quantization loss
Q8_0 33.2 GB High-precision quantization
F16 62.5 GB Half-precision weights

The model file accounts for only part of runtime memory. Applications also allocate the context cache, compute buffers, model metadata, and backend-specific overhead. CPU offloading can reduce GPU memory requirements, usually with lower generation speed.

Pro article

This story is for Pro members

You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.

Trending
  • No trending articles

Comments

avatar

Next Reads