Rakuten releases RakutenAI-3.0, 671B-parameter Japanese-optimized mixture-of-experts model
Rakuten Group has released RakutenAI-3.0, a 671 billion parameter mixture-of-experts (MoE) model designed specifically for Japanese language tasks. The model activates 37 billion parameters per token and supports a 128K context window. It is available under the Apache License 2.0 on Hugging Face.
Rakuten Releases 671B Parameter Model Optimized for Japanese
Rakuten Group has published RakutenAI-3.0, a 671 billion parameter mixture-of-experts language model engineered for Japanese language understanding and generation. The model activates 37 billion parameters per token and supports a 128,000 token context window.
Technical Specifications
The model uses a mixture-of-experts architecture, a design pattern that maintains computational efficiency by selectively activating only a subset of parameters for each input token. RakutenAI-3.0 is trained on a combination of publicly available open-source data and Rakuten's proprietary bilingual Japanese-English datasets.
Key specifications:
- Total parameters: 671 billion
- Active parameters per token: 37 billion
- Context window: 128,000 tokens
- Supported languages: Japanese and English
- Model format: F32, BF16, and F8_E4M3 quantization variants available
- License: Apache License 2.0
Deployment and Access
RakutenAI-3.0 is available on Hugging Face for download and local deployment. The company provides inference instructions using SGLang with recommended specifications requiring 8 tensor parallelism and 85% static memory allocation. The model has recorded 425 downloads in its first month on Hugging Face.
No official inference API or hosted endpoints have been announced. The model card indicates the model is not currently deployed by commercial inference providers.
Positioning
Rakuten positions RakutenAI-3.0 as delivering "superior grasp of Japanese language and culture" compared to existing models. The emphasis on Japanese-optimized training reflects increasing focus by regional technology companies on language-specific LLMs, following similar releases from companies like Alibaba (Qwen) and Baidu.
Limitations
Rakuten's documentation explicitly acknowledges that RakutenAI-3.0 can generate biased, inaccurate, or unsafe outputs like other large language models. The company recommends implementing appropriate safeguards for production deployments.
What This Means
Rakuten's entry into open-source Japanese-optimized LLMs signals sustained competition in regional language models. At 671B parameters with a 128K context window, it competes in scale with existing open models but targets a specific linguistic niche. The Apache 2.0 license and community release suggest Rakuten is prioritizing ecosystem participation over proprietary monetization, similar to Meta's approach with Llama. The model's availability only through local deployment (no hosted API) limits accessibility for developers without substantial compute resources.
Related Articles
Upstage Releases Solar Mini 4: 35B MoE Model with 524K Context at $0.05/$0.20 per Million Tokens
Upstage has released Solar Mini 4, a compact mixture-of-experts model with 35B total parameters, 3B active parameters, and a 524K token context window. The model targets agentic workloads and is priced at $0.05 per 1M input tokens and $0.20 per 1M output tokens, a promotional 50% discount off standard rates.
Perceptron Launches Mk1.5, a Multimodal Perception Model for Physical Agents with Structured Spatial Outputs
Perceptron has released Mk1.5, a perception model built for physical agents that accepts text, image, video, and audio input and returns text alongside structured spatial annotations. It succeeds Mk1 and is priced at $0.15 per 1M input tokens and $1.50 per 1M output tokens.
Black Forest Labs Releases FLUX 3 Action, a 7B Open-Weights World Action Model, Claims Top RoboLab Benchmark Score
Black Forest Labs has released FLUX 3 Action, a 7B parameter open-weights World Action Model. The company claims it achieves first place on the RoboLab benchmark, though independent verification is pending.
Black Forest Labs Releases FLUX 3 Action, a 7B-Parameter Open Robotics Model
Black Forest Labs has released FLUX 3 Action, an open-weight robotics model built on its FLUX 3 multimodal foundation. The 7-billion-parameter model reads multi-camera video feeds and predicts what a robot should do next, claiming a record success rate on the RoboLab-120 leaderboard while running nearly 4x faster than the previous best open model.
Comments
Loading...