Upstage Releases Solar Mini 4: 35B MoE Model with 524K Context at $0.05/$0.20 per Million Tokens
Upstage has released Solar Mini 4, a compact mixture-of-experts model with 35B total parameters, 3B active parameters, and a 524K token context window. The model targets agentic workloads and is priced at $0.05 per 1M input tokens and $0.20 per 1M output tokens, a promotional 50% discount off standard rates.
Solar Mini 4 — Quick Specs
Upstage Releases Solar Mini 4
Upstage has released Solar Mini 4, a compact mixture-of-experts (MoE) language model built for cost-sensitive, agentic workloads. The model is now available through OpenRouter and Upstage's API.
Specifications
Solar Mini 4 uses a 35-billion-parameter MoE architecture with only 3 billion parameters active per forward pass, a design intended to keep inference costs and latency low while retaining broad model capacity. It ships with a 524,288-token (524K) context window, positioning it for long-document and long-horizon agentic tasks.
According to Upstage, the model offers fluent Korean alongside strong English and Japanese performance, continuing the company's focus on multilingual support centered on Korean-language use cases.
Pricing
Solar Mini 4 is listed at $0.05 per 1 million input tokens and $0.20 per 1 million output tokens on OpenRouter — a promotional 50% discount off the standard listed rate of $0.10 / $0.40 per 1 million tokens. Cache-read pricing is listed at $0.01 per 1M tokens (input) and $0.005 per 1M tokens (output) under the discounted tier.
OpenRouter data shows throughput up to 95 tokens per second and P50 latency as low as 0.28 seconds across providers, with 100% uptime recorded over the most recent 24-hour and 3-day windows.
Deployment and Access
Upstage supports bring-your-own-key (BYOK) access with zero-data-retention (ZDR) enforcement. Organizations wanting ZDR enabled on their Upstage Console API key must contact Upstage directly to activate it; without this, standard data-retention terms apply.
Model Family Context
Solar Mini 4 joins two other current Upstage releases available on OpenRouter: Solar Pro 4, which shares the same 524K context window and is priced at $0.09 / $0.36 per 1M tokens for office- and coding-oriented workflows, and Solar Pro 3, a larger 102B-parameter MoE model (12B active) with a 131K context window priced at $0.15 / $0.60 per 1M tokens. Solar Mini 4 sits below both in parameter count and price, targeting lighter-weight, higher-volume use cases.
No benchmark scores (e.g., MMLU, HumanEval) or a training data cutoff date have been disclosed by Upstage for this release.
What This Means
Solar Mini 4's combination of a 524K context window and 3B active parameters puts it in a niche aimed squarely at high-throughput agentic pipelines — think document retrieval, multi-step tool use, and long-running conversational agents — where per-token cost and latency matter more than raw benchmark ceiling. The MoE design (35B total, 3B active) is a common strategy for keeping inference costs down while preserving some of the capacity benefits of a larger dense model, though without published benchmark scores it's not possible to independently verify how it compares to similarly priced competitors like smaller Gemini, Llama, or Qwen variants. Upstage's continued emphasis on Korean-language fluency, paired with English and Japanese support, signals the company is prioritizing regional strength over head-to-head competition with frontier labs on general-purpose leaderboards. The ZDR/BYOK requirement also suggests enterprise and compliance-focused customers are a target segment.
Related Articles
Xiaomi Releases MiMo-V2.6-Pro-RL, a 1.02T-Parameter Omnimodal Model with 1M-Token Context
Xiaomi's MiMo team has released MiMo-V2.6-Pro-RL, a 1.02-trillion-parameter sparse mixture-of-experts model with 42B active parameters, 1M-token context, and native text/image/video/audio processing. The model was trained via a single mixed reinforcement learning run spanning coding, agentic, visual, and cybersecurity tasks, with benchmark scores that Xiaomi claims approach or match Claude Opus 5 and GPT-5.6 on several agentic and coding tests.
Xiaomi Launches MiMo-V2.6-Pro-UltraSpeed: Same Quality, 10x Faster Output
Xiaomi's MiMo-V2.6-Pro-UltraSpeed is a fast-inference edition of the company's 1T-parameter flagship MiMo-V2.6-Pro, delivering roughly 10x the output speed at matching quality. It retains the 1M-token context window and native multimodal capabilities, priced at $4.35/$8.70 per 1M input/output tokens.
Anthropic Ships Claude Opus 5.5, OpenAI Launches GPT-6 Sol and Luna — All Cheaper Than Predecessors
Anthropic released Claude Opus 5.5 at $4/$20 per million input/output tokens, undercutting Opus 5's $5/$25 pricing while claiming better agentic coding scores. OpenAI countered with GPT-6 Sol ($2/$10) and GPT-6 Luna ($0.10/$0.50), both up to 50% cheaper than GPT-5.6's promotional rates.
OpenAI Launches GPT-6 Luna: Fast, Low-Cost Model With 1.1M Context Window
OpenAI has released GPT-6 Luna, the fast and cost-efficient entry in its new GPT-6 model family, featuring a 1.1M token context window and pricing starting at $0.10 per 1M input tokens. The model is positioned below GPT-6 Sol and GPT-6 Astra in OpenAI's tiered lineup.
Comments
Loading...