Cohere's North Small Translate Beats DeepL Across 50 Languages With Downloadable Weights

Cohere's open-weights 218B mixture-of-experts translation model tops WMT26 under 1T parameters, with 25B active params and no reasoning traces.

·
·
·
  • Cohere released North Small Translate, a 218B / 25B-active MoE translation model with open weights.
  • Scores 83.60 on WMT26 across 50+ languages, beating DeepL NextGen (81.37) and Google Translate (68.20).
  • Agentic multi-pass variant reaches 84.36 by detecting and fixing its own errors.
  • Non-reasoning design prioritizes throughput; runs on 2x H100 or 1x B200.
  • Trained with difficulty sampling plus a 5-step SFT, DPO, and online RL protocol with RWS.
  • Free on Cohere API until rate limits; CC BY-NC 4.0 weights, commercial license via Model Vault.

Cohere releases North Small Translate with downloadable weights

Cohere has released North Small Translate 1.0, a text-only machine-translation model with downloadable weights, support for more than 50 languages, and an optional multi-pass translation workflow. Developed with RWS, the sparse mixture-of-experts transformer contains 218 billion parameters while activating 25 billion for each token. Cohere recommends two H100 GPUs or one B200, allowing the quantized model to run on a single server.

North Small Translate targets high-volume localization, subtitling, support, and catalog workloads. The base model generates translations without long intermediate reasoning traces, reducing output tokens and inference latency. An optional agentic workflow spends additional inference time on multiple passes to improve quality. Cohere provides implementation details in the model documentation and evaluation results in its technical report.

Reported scores, with a vendor caveat

Cohere reports an all-language score of 83.60 on WMT26, a machine-translation benchmark released after the model’s training cutoff. That timing reduces the risk that benchmark examples appeared in the training data. The figures below come from Cohere’s evaluation and await independent replication; higher scores indicate better performance.

WMT26 all-language scores reported by Cohere
Model or system Score
North Small Translate with agentic workflow 84.36
North Small Translate 83.60
Qwen 3.5 397B A17B 81.56
DeepL NextGen 81.37
Gemma 4 31B 79.46
GLM 5.2 FP8 76.50
Google Translate 68.20

The single-pass model leads DeepL NextGen by 2.23 points and Qwen 3.5 397B A17B by 2.04 points in Cohere’s results. The 84.36 agentic score includes the multi-pass workflow and its additional inference cost. Qwen lists 397 billion total parameters and 17 billion active parameters, while North Small Translate uses 218 billion total and 25 billion active. Runtime comparisons therefore need to account for active parameters, numerical precision, sequence length, and hardware.

Cohere’s separate long-context evaluation translates two book chapters in one call and uses xCOMET-XL, an automated translation-quality metric, to score each paragraph. North Small Translate reached 48.9, compared with 21.3 for Google Translate and 19.4 for Gemma 4 31B. This evaluation uses a different scale from WMT26 and tests whether chapter-scale context helps preserve terminology, pronouns, and other cross-paragraph references.

Routing tokens, trimming latency

At inference time, a mixture-of-experts router sends each token through a subset of the network. North Small Translate selects eight of 128 routed experts for each token and also applies shared experts. This sparse design lowers per-token computation relative to activating all 218 billion parameters.

  • Routing: A sigmoid router scores the experts and normalizes the selected top-eight results.
  • Attention: Three sliding-window layers appear for each global-attention layer.
  • Local context: Sliding-window layers use a 4,096-token window with rotary positional embeddings.
  • Global context: Global layers omit positional embeddings.
  • Sequence limits: The model supports up to 16,000 input tokens and 16,000 output tokens.
  • Modalities: Inputs and outputs are text only.

The architecture adapts patterns from Cohere’s Command A family. Skipping generated reasoning traces reduces the number of decoded tokens, an important constraint for translation systems processing millions of segments. The optional Agentic Translation framework adds multiple passes for workloads that can absorb higher latency and compute costs.

Training on the stubborn examples

Cohere’s five-stage training protocol combines supervised fine-tuning, direct preference optimization, and online reinforcement learning. The team also used difficulty sampling to concentrate updates on documents the model still translated poorly. According to Cohere localization lead Tom Kocmi, an early model could already translate more than 90% of the training set, leaving many examples with little useful training signal.

Cohere builds harder, language-specific supervision with two methods that reduce dependence on a single teacher and convert post-edits into preference data:

  1. Best-per-Language Forward Translation: The strongest available system for each language generates target translations, allowing the teacher to vary by language.
  2. Post-Edit Driven Preference Distillation: Human or model post-edits become the preferred examples in direct-preference-optimization pairs, while the raw machine output serves as the rejected example.

Supervised fine-tuning also covers post-editing, translation-quality estimation, and general instruction following. Those capabilities support prompts that request error identification, terminology constraints, or revision of an existing translation alongside direct source-to-target translation.

Weights, API, and licensing

Cohere distributes an NVFP4 W4A16 checkpoint, which uses 4-bit weights and 16-bit activations. The weights carry the Creative Commons Attribution-NonCommercial 4.0 license, requiring attribution and limiting their use to non-commercial purposes under the license terms. Commercial access is available through Cohere Model Vault and RWS Language Weaver.

At launch, Cohere also offers North Small Translate through its API. Trial and production keys can use the model without usage charges until their applicable rate limits are reached. The model identifier is north-small-translate-1-0.

haskell
import os
from cohere import ClientV2

co = ClientV2(api_key=os.environ["COHERE_API_KEY"])

response = co.chat(
    model="north-small-translate-1-0",
    messages=[
        {
            "role": "user",
            "content": (
                "Translate into French:\n"
                "Enterprises need accurate translations."
            ),
        }
    ],
)

print(response.message.content[0].text)

Broad regional gains, defined limits

Across Cohere’s evaluation of 32 high-resource languages and 18 additional languages, North Small Translate showed less regional score variation than similarly sized competitors. Cohere reports that it outperformed DeepL NextGen across the Middle East and North Africa, South Asia, Southeast Asia, and East Asia. Its largest regional leads were approximately eight to 10 points in South Asia and the Middle East and North Africa.

  • Text only: Speech, images, scanned documents, and page layout require separate transcription, OCR, or document-processing systems.
  • Finite context: The 16K input limit requires chunking for full books and long documents.
  • Agentic overhead: The multi-pass workflow improves Cohere’s reported score while increasing latency and compute consumption.
  • Commercial restrictions: The downloadable CC BY-NC 4.0 weights do not permit unrestricted commercial deployment.

A practical self-hosted MT option

Downloadable weights give localization teams direct control over data placement, batching, terminology prompts, model serving, and infrastructure costs. North Small Translate also provides a specialized alternative to general-purpose language models whose translation ability comes with larger serving footprints or reasoning-oriented decoding behavior.

A production evaluation should cover representative language pairs, document lengths, named entities, terminology consistency, human quality ratings, throughput, peak GPU memory, and total serving cost. The single-pass model fits bulk translation, while selected difficult or high-value segments can use the agentic workflow when its quality gain justifies the extra inference time.

Trending
  • No trending articles

Comments

avatar

Next Reads