Kurdish Speech Logo
Kurdish Speech
← Back to articles
Cohere Releases North Small Translate: A 218B MoE Translation Model That Scores 83.6 on WMT26 Across 50 Languages
Machine Translation & Multilingual NLP

Cohere Releases North Small Translate: A 218B MoE Translation Model That Scores 83.6 on WMT26 Across 50 Languages

Cohere has released North Small Translate , an open-weight machine translation model from Cohere and Cohere Labs. It is a sparse Mixture-of-Experts (MoE) model with 218B total and 25B active parameters. It covers 50 languages, from Albanian to Vietnamese. On Cohere’s WMT26 evaluation, it scores 83.6 averaged across all languages. Cohere says that beats DeepL and Google Translate, plus open options like GLM 5.2 and Mistral Large 3. Is it deployable? Yes. Call it free on Cohere’s API until rate limits, self-host it non-commercially, or license it commercially. Back to Where the Transformer Started Google researchers introduced the Transformer in 2017 with Attention Is All You Need . Its main results came from WMT 2014 English-to-German and English-to-French translation. 9 years later, Cohere is returning to that original problem with a dedicated model. Cohere’s launch post on X frames translation as a sovereignty issue. Organizations that cannot communicate globally cannot stay sovereign. North Small Translate is the first translation model in Cohere’s North family. It follows Tiny Aya and Command A Translate in Cohere’s multilingual lineage. Cohere built it with RWS , whose Language Weaver scientists and language experts shaped its real-world quality. Architecture The model structure describes a decoder-only sparse MoE Transformer. Here are the key details: Experts: 128 experts, 8 activated per token, plus shared experts applied to every token. Router: A sigmoid over expert logits, normalized over the selected top-k. Attention: Sliding-window layers (window 4096, RoPE) and g

Source: MarkTechPost

Source: MarkTechPost