StepFun launches Step 3.7 Flash: 196B MoE model with 256K context and adjustable reasoning levels at $0.20/$1.15 per 1M
StepFun has released Step 3.7 Flash, a 196B-parameter Mixture-of-Experts model that activates approximately 11B parameters per token. The multimodal model supports a 256K context window and introduces selectable reasoning levels (high/medium/low), priced at $0.20 per 1M input tokens and $1.15 per 1M output tokens.
Step-3.7-Flash — Quick Specs
StepFun Launches Step 3.7 Flash with Adjustable Reasoning Levels
StepFun has released Step 3.7 Flash, a 196B-parameter Mixture-of-Experts (MoE) model that activates roughly 11B parameters per token during inference. The model includes native image and video understanding capabilities through an integrated vision encoder.
Technical Specifications
Step 3.7 Flash supports a 256K token context window and is priced at $0.20 per 1M input tokens and $1.15 per 1M output tokens. The model was released on May 28, 2025, according to OpenRouter's listing.
The architecture combines a 196B-parameter language backbone with a vision encoder, making it StepFun's latest multimodal offering. By activating only 11B parameters per token through its MoE design, the model aims to balance performance with computational efficiency.
Selectable Reasoning Levels
A distinctive feature is the model's three selectable reasoning levels—high, medium, and low—allowing developers to trade off between processing speed, cost, and reasoning depth based on specific use cases. This gives callers direct control over how the model allocates compute resources per query.
Target Use Cases
According to StepFun, Step 3.7 Flash is designed for coding tasks, agentic workflows, structured output generation, and long-context productivity applications. The 256K context window positions it for document analysis, extended code review, and multi-turn conversations requiring substantial memory.
The model is currently available through OpenRouter, which routes requests across multiple providers to handle different prompt sizes and parameters.
What This Means
Step 3.7 Flash represents StepFun's entry into the competitive space of large-context multimodal models, directly competing with offerings from Anthropic, Google, and others in the 200K+ context range. The adjustable reasoning levels are a notable differentiation—most models offer fixed inference patterns, while this approach lets developers optimize for their specific latency and quality requirements. The $0.20/$1.15 pricing puts it in the mid-tier range, though real-world performance benchmarks will determine whether the selectable reasoning modes deliver meaningful value beyond standard inference optimization.
Related Articles
Mistral AI Releases Shieldstral-1.0-3B, a 3B-Parameter Policy-Adaptive Safety Classifier
Mistral AI has released Shieldstral-1.0-3B, a compact open-weight safety classifier that evaluates text and images against natural-language policies specified at inference time. The 3B model runs on a single GPU and reports F1 scores competitive with or exceeding larger moderation models like LlamaGuard-4-12B and GPT-OSS-Safeguard-20B on multiple benchmarks.
Mistral Releases Shieldstral, a 3B Open-Weights Safety Classifier That Matches Models 7x Its Size
Mistral has released Shieldstral, a 3B open-weights safety classifier that reframes content moderation as a policy-adaptive question-answering task. The model claims to match or outperform guard models up to 7x its size on text safety and multimodal benchmarks, and runs on a single 16GB GPU.
LG AI Research Releases K-EXAONE 2.0, a 750B-Parameter Open-Weight MoE Model with 262K Context
LG AI Research has released K-EXAONE 2.0, a 750-billion-parameter mixture-of-experts language model with 37B active parameters, a 262,144-token context window, and support for 10 languages. The model is open-weighted under Apache 2.0 and claims competitive results against Qwen3.5, GLM-5.1, and DeepSeek-V4 Pro on reasoning, coding, and long-context benchmarks.
Alibaba Unveils Qwen3.8-Max, a 2.4T-Parameter Open-Weight Model for Coding and Agentic Work
Alibaba announced Qwen3.8-Max, a 2.4T-parameter flagship model targeting coding and long-horizon agentic work, with open weights promised for next week alongside Qwen3.8-27B. The model posted strong third-party benchmark results, ranking #4 in Frontend Code Arena and matching Claude Opus 4.7 on the Vals Index at roughly 2.3x lower cost.
Comments
Loading...