model releaseKwaipilot

Kwaipilot Releases KAT-Coder-Pro V2.5 with 256K Context Window at $0.74/$2.96 Per Million Tokens

TL;DR

Kwaipilot has released KAT-Coder-Pro V2.5, a coding-focused language model with a 256,000-token context window. The model is priced at $0.74 per million input tokens and $2.96 per million output tokens, available through OpenRouter.

1 min read
0

KAT-Coder-Pro V2.5 — Quick Specs

Context window256K tokens
Input$0.74/1M tokens
Output$2.96/1M tokens

Kwaipilot Releases KAT-Coder-Pro V2.5 with 256K Context Window

Kwaipilot has released KAT-Coder-Pro V2.5, a coding-focused language model with a 256,000-token context window. The model is priced at $0.74 per million input tokens and $2.96 per million output tokens.

The model is available exclusively through OpenRouter, which forwards requests directly to Kwaipilot without routing decisions. According to OpenRouter's data, prompt caching can reduce effective costs by 60-80% below list prices depending on the amount of repeated context.

Technical Specifications

  • Context Window: 256,000 tokens
  • Input Pricing: $0.74 per 1 million tokens
  • Output Pricing: $2.96 per 1 million tokens
  • Release Date: July 10, 2026
  • Model Type: Text (coding-focused)
  • API: OpenAI-compatible through OpenRouter

The model slug is kwaipilot/kat-coder-pro-v2.5 for API integration. OpenRouter provides monitoring data including throughput (tokens per second), latency, time-to-first-token (TTFT), and 30-day uptime metrics for the model.

Pricing Context

The output pricing of $2.96 per million tokens positions KAT-Coder-Pro V2.5 in the mid-range for coding models. The input cost of $0.74 per million tokens is competitive for models with 256K context windows.

OpenRouter's infrastructure automatically retries failed requests on alternative providers and tracks which applications send the most traffic to the model, providing signals about production workload patterns.

What This Means

The 256K context window allows developers to include substantial codebases or documentation in prompts without chunking. The pricing structure makes it viable for production coding assistance applications, though the 4:1 output-to-input price ratio means code generation workloads will cost significantly more than code analysis. The exclusive availability through OpenRouter rather than direct API access may limit adoption compared to models with multiple distribution channels.

Related Articles

model release

Tencent Releases Hy4 Preview: 770B-Parameter MoE Model with 1M Context for Coding Agents

Tencent has released Hy4 preview, a mixture-of-experts model with 770B total parameters and 49B active parameters, targeting coding agents and multi-step tool-use workflows. The model ships with a 1 million token context window and is priced at $0.834 per 1M input tokens and $2.501 per 1M output tokens.

model release

Tencent Open-Sources Hy4 Preview: 770B-Parameter MoE Model with 1M-Token Context

Tencent's Hy Team has open-sourced Hy4 preview, a 770-billion-parameter Mixture-of-Experts model with 49 billion activated parameters and a 1-million-token context window. The model is available under Apache 2.0 alongside an FP8-quantized variant, with Tencent claiming it beats GLM 5.3 and Kimi K3 on internal engineering evaluations.

model release

Alibaba Releases Qwen3.8 Flash, a Multimodal Reasoning Model with 1M-Token Context

Alibaba has released Qwen3.8 Flash, a multimodal reasoning model with a 1 million token context window, aimed at coding, agentic workflows, and visual/document analysis. It's priced at $0.16 per 1M input tokens and $0.47 per 1M output tokens through Alibaba Cloud International.

model release

Z.ai Confirmed as Creator of Chart-Topping 'Ox Alpha' Model, Weights Coming Wednesday

Z.ai, maker of the GLM model series, has confirmed it is behind Ox Alpha, the mysterious open-weight model that appeared anonymously on OpenRouter and topped benchmark leaderboards. The company will release the model's weights on Wednesday.

Comments

Loading...