model releaseTencent

Tencent Releases Hy-MT2-30B-A3B, a 30B-Parameter Translation Model with 3B Active Parameters

TL;DR

Tencent has released Hy-MT2-30B-A3B, a mixture-of-experts translation model with 30B total parameters and 3B active parameters, supporting 33 language pairs and five Chinese dialect and minority-language pairs. The model is available through Tencent Cloud at $0.074 per 1M input tokens and $0.295 per 1M output tokens.

2 min read
0

Hy-MT2-30B-A3B — Quick Specs

Context window8K tokens
Input$0.074/1M tokens
Output$0.295/1M tokens

Tencent has released Hy-MT2-30B-A3B, the flagship model in its Hy-MT2 translation family, according to a listing on OpenRouter. The model uses a mixture-of-experts architecture with 30 billion total parameters and 3 billion active parameters per inference pass, denoted by the "A3B" suffix.

What the model does

Hy-MT2-30B-A3B supports 33 language pairs alongside five Chinese dialect and minority-language pairs, according to Tencent. The model is built around five distinct translation workflows:

  • Structured translation — for documents with defined formatting
  • Delimiter-based translation — for segmented text
  • Contextual translation — accounting for surrounding text
  • Glossary-based translation — enforcing terminology consistency
  • Style-guided translation — matching a target tone or register

The model ships with an 8,000-token context window, which is modest compared to general-purpose large language models but in line with other dedicated translation systems that process shorter, discrete text segments rather than long documents in a single pass.

Pricing and availability

Hy-MT2-30B-A3B is available through Tencent Cloud at $0.074 per 1 million input tokens and $0.295 per 1 million output tokens. OpenRouter lists the model as newly added, with uptime and latency data still being collected — the platform notes there is "not enough uptime data to display yet" and no throughput figures are currently available. The listing shows a release date of August 20, 2026.

No independent benchmark scores were provided in Tencent's release materials or the OpenRouter listing. Tencent has not disclosed the training data composition, training cutoff date, or comparative evaluation results against other translation systems such as Google Translate's NMT models or DeepL's proprietary engines.

What this means

Hy-MT2-30B-A3B is a specialized tool, not a general-purpose chatbot competitor. By using a mixture-of-experts design with only 3B active parameters out of 30B total, Tencent is betting that translation-specific efficiency — cheap inference on a narrow task — matters more than raw scale for this use case. The pricing, at roughly a third of a cent per 1,000 output tokens, positions it as a low-cost option for high-volume translation pipelines, particularly for enterprises already inside the Tencent Cloud ecosystem.

The explicit support for Chinese dialects and minority languages signals a domestic-market focus, distinguishing it from Western translation APIs that prioritize major world languages. The lack of published benchmarks makes it difficult to independently verify translation quality against incumbents like DeepL or Google's translation stack, so buyers evaluating this model for production use should run their own comparative tests before committing volume. The 8K context ceiling also means it is best suited for sentence- or paragraph-level translation rather than long-document workflows without chunking.

Related Articles

model release

inclusionAI releases Ling 3.1 Flash: 560B MoE, 25B active, 262K context, free on OpenRouter

inclusionAI has released Ling 3.1 Flash, a hybrid reasoning mixture-of-experts model with 560B total and 25B active parameters and a 262K-token context window. It is listed as free on OpenRouter through NovitaAI. No benchmark scores have been published on the listing.

model release

China Telecom's Xing4.0-29B-A4B: 29B MoE, 4B Active, 256K Context, Trained Fully on Ascend NPUs

China Telecom AI's Xing4.0-29B-A4B (formerly the TeleChat line) is a mixture-of-experts model with 29B total and 4B active parameters and a native 256K context window, extensible to 512K. The company claims it is the first model of this scale trained entirely on Ascend NPUs with MindSpore. Community GGUF quantizations from Venastine-Research are already available.

model release

Cloudflare releases Clef decision models, claims 39 ms median latency vs. 524 ms for TypeSafe's Jev

Cloudflare has released Clef and Clef-flash, two open-weight decision models that return probabilities over predefined answer options instead of generating text. The company claims median latencies of 39 ms and 209 ms, against just over 524 ms for TypeSafe AI's Jev. Both support text and images and are API-compatible with Jev.

model release

Ai2 open-sources AstaBrief 8B, a Qwen3-8B report model it says runs 3.5x faster than Claude in Asta

Ai2 has open-sourced AstaBrief 8B, a model fine-tuned from Qwen3-8B that turns a research question and retrieved literature excerpts into a cited report. It is live in Asta as Fast mode, which averages 51.1 seconds per report versus 178.5 seconds for the Claude-powered Thinking mode, according to Ai2. The weights and training data are public.

Comments

Loading...