Johns Hopkins' Ettin Proves Smaller Encoders Beat Larger Decoders on Classification
Johns Hopkins and LightOn release ten paired encoder and decoder models trained identically, ending years of unfair architectural comparisons and beating ModernBERT, Llama 3.2, and SmolLM2.
PRO- Johns Hopkins and LightOn released Ettin, ten paired encoder and decoder models from 17M to 1B parameters.
- Both architectures share identical data, shape, and recipe, differing only in attention and training objective.
- Ettin encoders beat ModernBERT on GLUE and MTEB; decoders beat Llama 3.2 and SmolLM2 at matched sizes.
- A 400M encoder outperforms a 1B decoder on MNLI, with the reverse holding for generative tasks.
- Continued cross-objective pretraining for 50B tokens fails to close the architecture gap in either direction.
- Full release includes 236 checkpoints per model, batch-ordered training data, and the code repo.
For years, the encoder versus decoder debate has been a mess of confounding variables. When someone claimed a decoder-only model could replace BERT for classification, the models being compared had different training data, different parameter counts, different learning schedules, and different architectures. It was scientifically unsatisfying, and it left practitioners guessing about which family to reach for.
A team from Johns Hopkins and LightOn just fixed that with Ettin, named after the two-headed Norse giant. Ettin is the first suite of paired encoder-only and decoder-only models (17M-1B params) trained with identical data (2T tokens), architecture, and training recipes. The only things that differ between each pair are the attention pattern (bidirectional vs causal) and the training objective (masked vs causal language modeling).
Why apples-to-apples finally matters
Encoder-only models like BERT built the modern NLP era, but the field pivoted hard to decoder-only architectures because they generate text natively. That left a huge chunk of production workloads, classification, retrieval, embeddings, still leaning on models designed in 2019. Meanwhile, papers like LLM2Vec argued you could just take a 7B decoder, do a bit of continued pretraining with masked objectives, and get a better encoder for free.
Nobody could really test that claim rigorously, because a comparison with mismatched data and scale proves nothing. Ettin's authors built the controlled experiment the community has been missing.
The recipe, and why it beats ModernBERT
Ettin is essentially an open-data reproduction of the ModernBERT recipe, then applied symmetrically to decoders. The training data is a mix of DCLM combined with various curated sources from Dolma v1.7, drawing from the same sources used to train Olmo 2. Training proceeds in three phases:
This story is for Pro members
You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.