Microsoft's DeBERTa-v3-small Beats RoBERTa Using Half the Parameters

A compact 44M-parameter encoder from Microsoft Research keeps climbing the trending charts, offering ELECTRA-style efficiency and strong NLU accuracy at tiny inference cost.

·
·
·
Microsoft's DeBERTa-v3-small Beats RoBERTa Using Half the ParametersPRO
Read2 min
TypeModel
SubtopicSmall Models
  • DeBERTa-v3-small is a 44M-parameter English encoder with a 128K vocabulary from Microsoft Research.
  • Uses ELECTRA-style replaced-token-detection training instead of masked language modeling for better sample efficiency.
  • Introduces Gradient-Disentangled Embedding Sharing to kill the generator/discriminator tug-of-war on shared embeddings.
  • Scores 88.3 on MNLI-m and 82.8 F1 on SQuAD 2.0, matching models roughly twice its size.
  • Popular backbone for classification, NLI, prompt-injection detection, safety filters, and lightweight rerankers.
  • English-only and encoder-only; use mDeBERTa or decoder LLMs for generation or multilingual tasks.

Why Developers Keep Downloading DeBERTa-v3-small

Microsoft Research’s DeBERTa-v3-small drew about 777,000 Hugging Face downloads in the reported month. At the time of writing, the Hub also listed 213 fine-tunes and 20 quantized derivatives. Those changing figures measure artifact pulls, including automated jobs, but they show sustained activity around a model released before the current wave of decoder-only LLMs.

The appeal comes from a practical combination: competitive natural-language understanding scores, six transformer layers, broad framework support, and an MIT license. Its unusually large vocabulary complicates the size story, however. The transformer backbone has 44 million parameters, while the complete base checkpoint contains roughly 142 million.

The 44M label hides 142M weights

The model card’s 44 million figure covers the transformer backbone and excludes the token embedding table. A 128,000-token vocabulary with 768-dimensional embeddings adds about 98 million parameters, accounting for roughly 69% of the checkpoint.

DeBERTa-v3-small at a glance
Property Value
Transformer layers 6
Hidden size 768
Attention heads 12
Backbone parameters 44 million
Token embedding parameters About 98 million
Total base parameters About 142 million
Vocabulary 128,000 tokens
Maximum input length 512 tokens
Pretraining data 160 GB, following DeBERTa V2
License MIT

Half the backbone, competitive scores

Microsoft’s model card reports the following results. The parameter column covers each transformer backbone, keeping the comparison separate from vocabulary-dependent embedding sizes.

Benchmarks reported by Microsoft
Model Backbone parameters SQuAD 2.0 F1 / EM

Pro article

This story is for Pro members

You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.

Trending
  • No trending articles

Comments

avatar

Next Reads