coding models

8 articles tagged with coding models

July 20, 2026
benchmark

Moonshot's Kimi K3 tops Code Arena frontend benchmark at 1,679 points but scores only 39% on FrontierMath Tier 4

Moonshot AI's Kimi K3 model has claimed first place in the Code Arena frontend benchmark with a score of 1,679, surpassing Claude Fable 5 (1,631) and GPT-5.6 Sol (1,618). However, the model achieves only 39% accuracy on FrontierMath Tier 4, while top Western models from OpenAI and Anthropic reach near 90% on the same expert-level math tasks.

July 18, 2026
model release+1

Moonshot AI's Kimi K3 tops coding benchmarks, priced 50% below OpenAI's GPT-5.6 Sol

Beijing-based Moonshot AI released its Kimi K3 model Friday, which topped Arena's front-end coding capability rankings. The model is priced at half the cost of OpenAI's GPT-5.6 Sol, according to Bank of America research analysts, marking what Arena CEO calls "the single biggest release of the year."

July 9, 2026
model releaseOpenAI+1

OpenAI releases GPT-5.6 with three model variants, claims 80-point Coding Agent Index score for Sol

OpenAI released GPT-5.6 in three variants: Sol ($5 input/$30 output per 1M tokens), Terra ($2.50/$15), and Luna ($1/$6). According to OpenAI, Sol achieves an 80-point score on the Artificial Analysis Coding Agent Index, 2.8 points above Anthropic's Fable 5, while using less than half the output tokens and costing one-third less.

model release

Meta launches Muse Spark 1.1 coding model at $1.25/$4.25 per million tokens

Meta publicly released Muse Spark 1.1, a multimodal AI model designed for agentic coding workflows. The model is priced at $1.25 per million input tokens and $4.25 per million output tokens, positioning it slightly above Anthropic's Claude Haiku 4.5 and OpenAI's GPT-5.6 Luna.

July 8, 2026
model release

SpaceXAI launches Grok 4.5 at $2/$6 per million tokens, targets coding and enterprise work

Elon Musk's SpaceXAI has released Grok 4.5, priced at $2 per million input tokens and $6 per million output tokens. The model, trained alongside recently-acquired Cursor, is positioned as a coding and enterprise tool that claims to outperform Claude Opus 4.8 on several benchmarks while undercutting it on price by 60-76%.

May 20, 2026
model releasexAI

xAI Launches Grok Build 0.1: Coding Model with 256K Context for Agentic Workflows

xAI has released Grok Build 0.1, a coding-specialized model with a 256K context window and unlimited text output. The model is designed for agentic software engineering workflows and powers xAI's Grok Build CLI tool.

April 24, 2026
model releaseDeepSeek

DeepSeek Releases V4 Pro: 1.6T Parameter MoE Model with 1M Token Context at $1.74/M Input Tokens

DeepSeek has released V4 Pro, a Mixture-of-Experts model with 1.6 trillion total parameters and 49 billion activated parameters. The model supports a 1-million-token context window and costs $1.74 per million input tokens and $3.48 per million output tokens.

April 20, 2026
model releaseMoonshot AI+1

Moonshot AI Releases Kimi K2.6: 1T-Parameter MoE Model with 256K Context and Agent Swarm Capabilities

Moonshot AI has released Kimi K2.6, an open-source multimodal model with 1 trillion total parameters (32B activated) and 256K context window. The model achieves 80.2% on SWE-Bench Verified, 58.6% on SWE-Bench Pro, and supports horizontal scaling to 300 sub-agents executing 4,000 coordinated steps.