DeepSeek Drops 67B Open-Weight Model That Beats Llama2 on Coding and Math
DeepSeek's 67B base model lands on Hugging Face with full weights, commercial license, and plug-and-play support for vLLM, SGLang, and Transformers
PRO- Now on Hugging Face: DeepSeek LLM 67B Base weights are live at deepseek-ai/deepseek-llm-67b-base with ~8,650 downloads.
- Architecture: 67B parameter transformer with Grouped-Query Attention (GQA), trained on 2 trillion English and Chinese tokens.
- Strong benchmarks: Outperforms Llama2 70B Base on reasoning, coding, math, and Chinese; 73.78 HumanEval Pass@1 on the chat variant.
- Self-host only: No Hugging Face Inference Provider support; requires ~135 GB VRAM or quantized alternatives (GGUF, AWQ, GPTQ).
- Commercial use allowed under DeepSeek's model license; code repo is MIT licensed.
- Best for fine-tuning: Base model targets domain adaptation and research, not direct chat deployment.
DeepSeek LLM 67B Base is now available on Hugging Face, giving anyone with the right hardware direct access to the full model weights. The listing comes with out-of-the-box support for Transformers, vLLM, SGLang, and Docker, making it one of the more frictionless large open-weight models to self-host. The model has already pulled nearly 8,650 downloads and over 2,700 likes on the platform.
What it is
DeepSeek LLM 67B Base is a 67-billion-parameter language model trained from scratch on 2 trillion tokens in both English and Chinese. DeepSeek released both 7B and 67B sizes in base and chat variants, explicitly targeting the research community. The base variant published here is the raw pretrained model, before any instruction tuning, making it the right starting point for fine-tuning rather than direct chat use.
Under the hood
The model follows the same auto-regressive transformer decoder architecture as LLaMA. The 7B uses standard Multi-Head Attention, while the 67B uses Grouped-Query Attention (GQA). GQA is an efficiency technique that shares key and value projections across multiple query heads, reducing computational requirements while maintaining strong performance and shrinking the memory footprint compared to standard Multi-Head Attention at that scale.
The architecture also incorporates pre-norm structure and rotational position embeddings (RoPE), mixing in languages and tasks throughout training. Training used the AdamW optimizer with a sequence length of 4096. The 67B model was trained with a batch size of 4608 and a learning rate of 3.2e-4. A multi-step learning rate schedule was used, starting with 2000 warmup steps, then stepping down to 31.6% of the peak at 1.6 trillion tokens and 10% at 1.8 trillion tokens.
This story is for Pro members
You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.