Sakana AI's Namazu Beats Google Translate and DeepL on Japanese Cultural Nuance

Sakana AI upgraded its free JP-EN-ZH translator to the new Namazu model, beating Google Translate, DeepL, and Claude Opus 4.8 in head-to-head evaluations.

·
·
Sakana AI's Namazu Beats Google Translate and DeepL on Japanese Cultural Nuance
AuthorSakana AI
Read4 min
  • Sakana Translate now runs on the new-generation Namazu model, post-trained on Kimi K2.6 for Japanese context.
  • In 160 JP-EN tasks judged by TransEvalnia, it beat Google Translate 66.9%, Opus 4.8 56.8%, DeepL 54.0%.
  • Three modes retained: Translate up to 5,000 chars streaming, Proofread with diffs, Ask for nuance.
  • Bidirectional Japanese, English, Chinese translation remains free with an account at translate.sakana.ai.
  • Focus is culturally loaded expressions like Omiyamairi, honorifics, and business register that general MT misses.
  • Roadmap: file translation (PDF, Office), glossaries, API access, SSO, audit logs, on-prem deployment.

Sakana AI just refreshed the model behind its free translation service, swapping in the new generation of Sakana Namazu, the Japanese-tuned model series it launched earlier this year via API. The service still handles bidirectional translation between Japanese, English, and Chinese, but the output is meant to read more naturally, especially when the source text is loaded with Japanese cultural context.

Sakana Translate stays free for anyone with an account, and the same three modes carry over: Translate, Proofread, and Ask. What changed is the engine, and Sakana is making a specific claim about it: on their own head-to-head evaluations, the new model beats Google Translate, DeepL, and Claude Opus 4.8 on Japanese-to-English tasks.

Deep translation, not just word swapping

The update centers on a concept Sakana calls translating Japanese deeply. Rather than train a giant model from scratch, Sakana takes strong open-weight foundations and post-trains them for Japanese language and culture. The current generation of Namazu is built on Moonshot AI's open model Kimi K2.6 and further trained for Japanese and for Japanese business contexts.

The blog post uses Omiyamairi as the illustrative example, the traditional first shrine visit for a newborn. General translators tend to either transliterate it or produce something clunky. Sakana argues that the interesting failure mode of large general-purpose models is not grammar but culture, and that a smaller model tuned on the right nuances can beat a bigger one on this specific axis.

The evaluation numbers

To make that claim concrete, Sakana ran a head-to-head study using its own evaluator, TransEvalnia. TransEvalnia is a prompting-based translation evaluation and ranking system that uses reasoning in performing its evaluations and ranking, presenting fine-grained multidimensional evaluations based on a subset of the Multidimensional Quality Metrics. In plain terms, it asks an LLM judge to compare two translations across specific criteria like word choice appropriateness and naturalness, then explain its ranking.

The test set was 160 in-house Japanese-to-English tasks. Sakana Translate was judged better more than 50% of the time against every system in the comparison.

Bar chart showing Sakana Translate win rates against Google Translate, Claude Opus 4.8, and DeepL
  • vs Google Translate: 66.9% preferred
  • vs Claude Opus 4.8: 56.8% preferred
  • vs DeepL: 54.0% preferred

Two caveats worth flagging. First, this is an internal benchmark scored by an internal evaluator, not a public leaderboard like WMT. Second, the previous release of Sakana Translate reported an XCOMET-XL score of 0.835, while Google's Gemini 3.1 Pro scored 0.851 and OpenAI's GPT-5.5 scored 0.843, placing them at the top on a public benchmark. The story here is that Namazu is competitive on general MT metrics and pulls ahead specifically on Japanese cultural content.

What you actually get in the app

The three modes stay identical to the previous release, just running on the new weights:

  1. Translate: paste in text and get a streaming translation, up to roughly 5,000 characters at a time.
  2. Proofread: rewrites your draft into more natural phrasing, with changes shown as highlighted diffs covering grammar, tone, politeness, and register.
  3. Ask: follow-up questions on any translation or proofread in the same context, useful for probing why one honorific was chosen over another or requesting alternative phrasings.

When it makes sense to reach for it

If you are moving text between English and Japanese where cultural register matters, such as customer communications, marketing copy, or anything involving keigo (honorifics), this is worth benchmarking against your current setup. The stated strengths are the places general MT usually stumbles: business honorifics, cultural concepts, place names, proper nouns, and everyday context. For bulk technical translation or non-Japanese language pairs, the general-purpose frontier models are still the default choice.

What is coming next

Sakana laid out a roadmap that pushes the product toward enterprise use. Planned additions include file translation for PDFs and Microsoft Office documents, glossary integration, industry-specific tuned models, an API surface, and enterprise features like SSO, audit logs, and on-premises deployment. That direction suggests the free web app is functioning as a demo for a paid B2B service aimed at Japanese enterprises that need translation quality with data controls.

Comments

avatar