model releaseStepFun

StepFun launches Step 3.7 Flash: 196B MoE model with 256K context and adjustable reasoning levels at $0.20/$1.15 per 1M

TL;DR

StepFun has released Step 3.7 Flash, a 196B-parameter Mixture-of-Experts model that activates approximately 11B parameters per token. The multimodal model supports a 256K context window and introduces selectable reasoning levels (high/medium/low), priced at $0.20 per 1M input tokens and $1.15 per 1M output tokens.

2 min read
0

Step-3.7-Flash — Quick Specs

Context window256K tokens
Input$0.2/1M tokens
Output$1.15/1M tokens

StepFun Launches Step 3.7 Flash with Adjustable Reasoning Levels

StepFun has released Step 3.7 Flash, a 196B-parameter Mixture-of-Experts (MoE) model that activates roughly 11B parameters per token during inference. The model includes native image and video understanding capabilities through an integrated vision encoder.

Technical Specifications

Step 3.7 Flash supports a 256K token context window and is priced at $0.20 per 1M input tokens and $1.15 per 1M output tokens. The model was released on May 28, 2025, according to OpenRouter's listing.

The architecture combines a 196B-parameter language backbone with a vision encoder, making it StepFun's latest multimodal offering. By activating only 11B parameters per token through its MoE design, the model aims to balance performance with computational efficiency.

Selectable Reasoning Levels

A distinctive feature is the model's three selectable reasoning levels—high, medium, and low—allowing developers to trade off between processing speed, cost, and reasoning depth based on specific use cases. This gives callers direct control over how the model allocates compute resources per query.

Target Use Cases

According to StepFun, Step 3.7 Flash is designed for coding tasks, agentic workflows, structured output generation, and long-context productivity applications. The 256K context window positions it for document analysis, extended code review, and multi-turn conversations requiring substantial memory.

The model is currently available through OpenRouter, which routes requests across multiple providers to handle different prompt sizes and parameters.

What This Means

Step 3.7 Flash represents StepFun's entry into the competitive space of large-context multimodal models, directly competing with offerings from Anthropic, Google, and others in the 200K+ context range. The adjustable reasoning levels are a notable differentiation—most models offer fixed inference patterns, while this approach lets developers optimize for their specific latency and quality requirements. The $0.20/$1.15 pricing puts it in the mid-tier range, though real-world performance benchmarks will determine whether the selectable reasoning modes deliver meaningful value beyond standard inference optimization.

Related Articles

model release

GLM-5.3-Flash Debuts as Zhipu AI's First Natively Multimodal Model, 320B Parameters with 18B Active

Zhipu AI has released GLM-5.3-Flash, the first natively multimodal model in its GLM-5 series, built on a 320B-parameter mixture-of-experts architecture with only 18B active parameters. The company claims it outperforms GLM-5.2 while approaching Claude Opus 4.8 on coding and agentic benchmarks at a fraction of the cost. Unsloth has published quantized GGUF versions for local inference.

model release

Z.ai Launches GLM-5.3-Flash: 1M-Token Context, Image Support, Claimed 10x Cost Cut Over GLM-5.2

Z.ai has released GLM-5.3-Flash, a 320-billion-parameter Mixture-of-Experts model with 18 billion active parameters, a 1-million-token context window, and image input support. The model launched on LM Studio's Bionic platform hours after its official unveiling, with LM Studio claiming it is 9-10x cheaper to run than GLM-5.2.

model release

Tencent Open-Sources Hy4 Preview: 770B-Parameter MoE Model with 1M-Token Context

Tencent's Hy Team has open-sourced Hy4 preview, a 770-billion-parameter Mixture-of-Experts model with 49 billion activated parameters and a 1-million-token context window. The model is available under Apache 2.0 alongside an FP8-quantized variant, with Tencent claiming it beats GLM 5.3 and Kimi K3 on internal engineering evaluations.

model release

Tencent Releases Hy4 Preview: 770B-Parameter MoE Model with 1M Context for Coding Agents

Tencent has released Hy4 preview, a mixture-of-experts model with 770B total parameters and 49B active parameters, targeting coding agents and multi-step tool-use workflows. The model ships with a 1 million token context window and is priced at $0.834 per 1M input tokens and $2.501 per 1M output tokens.

Comments

Loading...