pricing
50 articles tagged with pricing
xAI Releases Grok 4.7 at $2/$6 per Million Tokens, Trails Claude and GPT-6 on Benchmarks
xAI has launched Grok 4.7 at $2 per million input tokens and $6 per million output tokens, undercutting Western rivals on price. But independent benchmarks show it trailing Claude Fable 5.1 and GPT-6 by a wide margin, especially in agentic coding.
Qwen3.8-Omni-Flash Prices Multimodal AI at $0.15/$0.47 per Million Tokens, Undercutting Gemini Flash by 5x
Alibaba's Qwen team released Qwen3.8-Omni-Flash, a multimodal model for AI agents that processes audio and video with a 1 million token context window. Pricing undercuts Google's Gemini 3.8 Flash by roughly 5x on input and 8x on output, according to Qwen.
OpenRouter Adds DeepSeek Flash Latest Alias With 1M-Token Context Window
OpenRouter has launched deepseek-flash-latest, a persistent endpoint that always points to the current DeepSeek Flash model. It offers a 1,049K token context window, text-and-image input, and pricing of $0.15 per 1M input tokens and $0.60 per 1M output tokens.
OpenAI Launches GPT-6 Astra With Half the Message Allowance of GPT-5.6 Sol
OpenAI has begun rolling out GPT-6 Astra to top-tier ChatGPT plans, the API, Azure, and AWS Bedrock. The model delivers roughly half the usage allowance of GPT-5.6 Sol across comparable plans, with Plus and Business users gaining access in the coming days.
OpenAI Lists GPT-6 Astra Pro on OpenRouter: Same Model, Higher-Compute Reasoning Mode
GPT-6 Astra Pro, now listed on OpenRouter, is the existing GPT-6 Astra model configured to run with reasoning.mode set to 'pro' for higher-quality output on complex tasks. It carries a 1M-token context window and tiered pricing from $5/$25 to $20/$100 per million input/output tokens depending on the serving tier.
xAI Brings Grok Bot to iPad and Android, Cuts Price From $300/Month to $20/Month
xAI has expanded its Grok Bot AI agent app from iPhone and Mac to iPad and Android, while slashing the price from $300/month to compatibility with the $20/month Cursor Pro plan. Enterprise customers using Grok or Cursor get free access for a limited time.
OpenAI's GPT-6 Astra Reportedly Automates AI Engineering Tasks at Under $6 an Hour, According to Latent Space Testing
A Latent Space report describes GPT-6 Astra, a new OpenAI model the blog says can autonomously handle AI engineering tasks—training models, labeling data, deploying systems—at an estimated cost of under $6 per hour. The claims, including 97.6% on FrontierMath and 99.9% on ARC-AGI-3, come from independent blog testing rather than an official OpenAI announcement.
Meta Releases Muse Spark 1.3, Cheapest Model in Its Performance Class at $0.55 Per Task
Meta has released Muse Spark 1.3, its fourth model in five months, with an xhigh tier available now and a more powerful max tier in limited preview. The model improves sharply on agentic benchmarks and costs $0.55 per index task—cheaper than any rival at the same performance level—but still trails Claude Fable 5.1 on most tests.
Google Launches Gemini 3.8 Flash, Warns It May Use More Tokens Despite Unchanged Pricing
Google released Gemini 3.8 Flash just weeks after Gemini 3.7 Flash, keeping the same per-token pricing of $0.75/$3.75 per million input/output tokens but warning it may consume more tokens overall. The model also ships with a cyber-focused variant restricted to a new government partner program called Fairwind.
Google Rolls Out Gemini 3.8 Flash, Third Flash Update in Three Months
Google has released Gemini 3.8 Flash, the third Flash-tier update in three months, arriving just three weeks after Gemini 3.7 Flash. The model is live now in the Gemini app, Google Antigravity, and AI Studio with introductory pricing of $0.75/1M input and $3.75/1M output tokens.
Anthropic Releases Claude Fable 5.1 and Mythos 5.1, Cuts Cache Pricing 75% But Output Tokens Jump 70%
Anthropic launched Claude Fable 5.1 and Claude Mythos 5.1, claiming the top spot on Artificial Analysis's Intelligence Index at 66. Cache-read pricing dropped 75% to $0.25 per million tokens, but a 1.7x increase in output token usage pushes net per-task cost up 20%.
OpenRouter Adds Auto-Updating Alias for Zhipu AI's GLM Flash Model Family
Z.ai has published GLM Flash Latest on OpenRouter, a routing alias that automatically points to the newest checkpoint in the GLM Flash lineup. It supports a 1.31M token context window and multimodal text, image, and video input at $0.07 per 1M input tokens and $0.25 per 1M output tokens.
Anthropic Releases Claude Fable 5.1, Claims 52.6% on New Terminal-Bench-Science Benchmark
Anthropic released Claude Fable (and Mythos) 5.1, claiming a 52.6% score on the new Terminal-Bench-Science 0.1 benchmark — up sharply from 24.7% for Fable 5. Independent testing shows the model's five reasoning levels produce dramatically different output token counts and costs for identical prompts, ranging from $0.10 to $3.30 per request.
Anthropic Releases Claude Fable 5.1, Cuts Agentic Workload Pricing Up to 45%
Anthropic has released Claude Fable 5.1, an upgrade to its top-tier Fable 5 model launched in June, alongside a restricted-access sibling called Mythos 5.1. The company claims the new model matches or beats Fable 5's performance while cutting costs by up to 45% on agentic workloads through reduced cache-read pricing.
OpenAI Quietly Rolls Out Outcome-Based Pricing, Charging Some Customers Only When Tasks Succeed
OpenAI has quietly begun offering some large customers a pay-per-outcome model, charging only when its AI successfully completes tasks such as customer support requests, according to The Information. The shift joins a broader industry move away from flat subscriptions toward usage- and results-based billing, led by startups like Sierra, Fin, and Cognition.
Anthropic to Cut Claude Code Weekly Limits by 17% Despite Calling It a 25% Increase
Anthropic will permanently raise Claude Code's baseline weekly usage limits by 25% starting September 14. Because this replaces a temporary 50% boost currently active, users will actually end up with about 17% less capacity than they have today.
Google Releases Gemini 3.7 Flash, Cuts Price in Half Versus 3.6 Flash
Google has released Gemini 3.7 Flash, just three weeks after Gemini 3.6 Flash, claiming substantial gains in coding, web development, and document reasoning. The model launches at an introductory price of $0.75 per 1M input tokens and $3.75 per 1M output tokens — half the cost of its predecessor.
xAI's Grok 4.6 Matches Claude and GPT-5.6 on Benchmarks, Costs 60% Less
xAI's Grok 4.6 ties OpenAI's GPT-5.6 Sol on the Artificial Analysis Intelligence Index with a score of 61, trailing only Anthropic's Claude Opus 5 and Claude Fable 5. Pricing remains at $2/$6 per million tokens, undercutting both competitors by more than 60 percent.
Microsoft's MAI Code 1.1 Flash Loses on Price and Performance to DeepSeek-V4-Flash
Microsoft's MAI Code 1.1 Flash beats its predecessor and mini-models from Anthropic and OpenAI on SWE-bench Verified, but DeepSeek-V4-Flash-0731 outperforms it on Terminal Bench 2.1 (82.7% vs 62.9%) while costing roughly a third as much per token.
OpenAI Adds $125/Month Premium Seats to ChatGPT Business for Heavy Agentic Use
OpenAI is introducing Premium Seats for ChatGPT Business at $125 per user per month ($100 with annual billing), offering five times the usage capacity of standard seats and removing the five-hour usage limit. Standard seats remain unchanged at $25 per month.
OpenAI Launches $125/Month ChatGPT Business Premium Seat With 5x Usage Limits
OpenAI has launched ChatGPT Business Premium seats, a new tier priced at $125/month ($100 if billed annually) that offers five times the usage of Standard Business seats and removes the five-hour-per-day limit on advanced features. The move comes as Chinese open-weight models increasingly rival closed-source frontier AI on capability.
OpenAI Removes Text Chat Limits for ChatGPT Free and Go Users, Upgrades GPT-5.6 Sol for Plus and Pro
OpenAI will remove text chat rate limits for ChatGPT Free and Go users starting next week and add a 'Think' button for deeper reasoning. Plus and Pro subscribers get an updated GPT-5.6 Sol model that OpenAI claims is more accurate with facts, dates, and sourcing.
DeepSeek Launches 'V4 Flash Latest' Alias with 1M+ Token Context on OpenRouter
DeepSeek has published a new routing endpoint, deepseek-v4-flash-latest, that always points to the newest model in its V4 Flash family. The endpoint offers a 1,049K token context window and pricing of $0.09/M input and $0.18/M output tokens via OpenRouter.
DeepSeek V4-Flash 0731 Update Jumps Terminal-Bench Score by 25.8 Points With No Architecture Change
DeepSeek released V4-Flash 0731, a post-training-only update to its API and open-weights model that lifted Terminal-Bench scores by 25.8 points without changing model architecture or parameter count. The update arrived alongside disclosed sandbox-escape incidents at OpenAI and Anthropic that renewed debate over eval infrastructure and open-weight safety.
DeepSeek V4 Flash 'O731' Nearly Matches GPT-5.6 Luna, Costs 60% Less to Run
DeepSeek has updated its budget model V4 Flash to version '0731,' pushing its Artificial Analysis Intelligence Index score to 50 — just one point behind OpenAI's GPT-5.6 Luna — while costing an estimated 60 percent less per task. The MIT-licensed model keeps its 284B-parameter architecture but shows major gains in agentic benchmarks and token efficiency.
OpenAI Cuts GPT-5.6 Prices Up to 80%, Says Model's Own Self-Optimization Work Drove the Savings
OpenAI cut GPT-5.6 Luna pricing by 80% to $0.20/$1.20 per million tokens and GPT-5.6 Terra by 20% to $2/$12, while adding a 2.5x-faster mode for Sol at double the price. The company says GPT-5.6 itself rewrote production inference kernels and tuned its own speculative decoding pipeline to enable the cuts.
OpenAI Slashes GPT-5.6 Luna Pricing by 80%, Cuts Terra by 20%
OpenAI cut GPT-5.6 Luna pricing by 80 percent to $0.20 per million input tokens and $1.20 per million output tokens, while Terra dropped 20 percent to $2/$12. The company attributes the cuts to infrastructure efficiency gains and mounting price competition, particularly from Chinese providers.
OpenAI Cuts GPT-5.6 Terra Price 20%, Luna Price 80% Across API and ChatGPT
OpenAI is cutting API prices for its GPT-5.6 Terra and Luna models by 20% and 80%, respectively, compared to prices set earlier this month. The company says the lower costs are also reflected in usage limits for ChatGPT Work and Codex subscribers, though subscription prices remain unchanged.
OpenAI Cuts GPT-5.6 Luna Price 80%, Terra 20%, as Enterprise Cost Pressure Mounts
OpenAI is cutting the price of GPT-5.6 Luna by 80% and GPT-5.6 Terra by 20%, just three weeks after launching the models. The move comes as enterprises grow more cost-conscious and rivals including Anthropic, Google, and Moonshot AI push cheaper alternatives.
Cursor Launches ₹649/Month 'Start' Plan for Indian Developers
Cursor has launched Cursor Start, a new ₹649 per month subscription plan tailored for developers in India, complete with local UPI billing. The plan includes access to Grok 4.5, Cursor's Composer model, always-on cloud agents, and Cursor for iOS.
Anthropic Launches Claude Opus 5 (Fast) at $10/$50 per Million Tokens, 1M Context Window
Anthropic has released Claude Opus 5 (Fast), a higher-throughput variant of Opus 5 that carries identical capabilities but runs at roughly 2x the price of the standard model. The model ships with a 1 million token context window and is available now through OpenRouter.
Moonshot AI Releases Kimi K3: 2.8T Parameter Open Model at $3/$15 Per Million Tokens
Moonshot AI has released Kimi K3, a 2.8 trillion parameter model with 1 million token context window and native multimodal input. The model ranks #1 in Frontend Code Arena and #9 in Text Arena, with pricing at $3 per million input tokens and $15 per million output tokens—comparable to Claude Sonnet 5 pricing while delivering performance the company claims is near Claude Opus 4.8 and GPT-5.5.
Anthropic reverses course, makes Claude Fable 5 permanent on subscription plans
Anthropic announced July 18 that Claude Fable 5 will remain available on subscription plans, reversing its previous decision to make the model API-only. Max and Team Premium subscribers will receive access at 50% of standard limits starting July 20, while Pro and Team Standard users get a one-time $100 credit.
Moonshot AI releases 2.8T parameter Kimi K3, pricing at $3/$15 per million tokens
Chinese AI lab Moonshot AI released Kimi K3, a 2.8 trillion parameter model priced at $3 per million input tokens and $15 per million output tokens. The model is currently available via API, with open weights promised by July 27, 2026. This represents the most expensive pricing from a Chinese AI lab to date, matching Anthropic's Claude Sonnet series.
Anthropic launches rupee pricing for Claude in India at ₹2,000/month, its second-largest market
Anthropic has begun displaying rupee-denominated pricing for Claude subscriptions in India, its second-largest market after the US with 5.8% of global usage. Claude Pro is priced at ₹2,000 ($21) monthly when billed annually, compared to $17 in the US, with Indian prices including local taxes.
Anthropic extends Claude Fable 5 access through July 19 amid GPT-5.6 Sol competition
Anthropic has extended Claude Fable 5 access on all paid plans through July 19, 2026, marking another extension of the advanced model's availability. The extension comes after OpenAI released GPT-5.6 Sol, which is classified in the same Fable/Mythos model tier.
OpenAI releases GPT-5.6 family in three sizes: Luna at $1/$6, Terra at $2.50/$15, Sol at $5/$30 per 1M tokens
OpenAI released its GPT-5.6 flagship model family in three sizes: Luna ($1/$6 per 1M tokens), Terra ($2.50/$15), and Sol ($5/$30). The company claims GPT-5.6 Sol scores 53.6 on the Agents' Last Exam benchmark, outperforming Claude Fable 5's score by 13.1 points.
Meta launches Muse Spark 1.1 coding model at $1.25/$4.25 per million tokens
Meta publicly released Muse Spark 1.1, a multimodal AI model designed for agentic coding workflows. The model is priced at $1.25 per million input tokens and $4.25 per million output tokens, positioning it slightly above Anthropic's Claude Haiku 4.5 and OpenAI's GPT-5.6 Luna.
OpenAI Releases GPT-5.6 Luna Pro with Extended Reasoning Mode at $1/$6 Per Million Tokens
OpenAI has released GPT-5.6 Luna Pro, a reasoning-enhanced variant of GPT-5.6 Luna with a 1 million token context window. The model is priced at $1 per million input tokens and $6 per million output tokens, with a knowledge cutoff date of February 2026.
OpenAI Releases GPT-5.6 Sol Pro with Extended Reasoning Mode at $5 Input/$30 Output per 1M Tokens
OpenAI has released GPT-5.6 Sol Pro, a reasoning-enhanced variant of GPT-5.6 Sol designed for complex tasks. The model features a 1 million token context window, February 2026 knowledge cutoff, and is priced at $5 per 1M input tokens and $30 per 1M output tokens.
OpenAI Releases GPT-5.6 Luna: $1/$6 Per 1M Tokens With 1M Context Window
OpenAI has released GPT-5.6 Luna, a fast and cost-efficient model in its GPT-5.6 series. The model features a 1 million token context window and is priced at $1 per 1M input tokens and $6 per 1M output tokens, with a knowledge cutoff of February 2026.
OpenAI Releases GPT-5.6 Terra: Mid-Tier Model at $2.50 Input/$15 Output per 1M Tokens
OpenAI has released GPT-5.6 Terra, a mid-tier model in its GPT-5.6 series priced at $2.50 per million input tokens and $15 per million output tokens. The model features a 1 million token context window and February 2026 knowledge cutoff, positioned between the flagship Sol and cost-efficient Luna tiers.
SpaceXAI releases Grok 4.5 at $2/$6 per million tokens, trained with Cursor on NVIDIA GB300 GPUs
SpaceXAI has released Grok 4.5, the first model developed after its rebrand from xAI and trained in partnership with Cursor. The model is priced at $2 per million input tokens and $6 per million output tokens, and was trained across tens of thousands of NVIDIA GB300 GPUs focused on coding, science, engineering, and math datasets.
xAI releases Grok 4.5 at $2/$6 per million tokens, claims Opus 4.7 performance at 60% lower cost
xAI has released Grok 4.5, pricing it at $2 per million input tokens and $6 per million output tokens — significantly undercutting Anthropic's Opus 4.7 ($5/$25 per million). Elon Musk claims the model delivers comparable performance to Opus 4.7 while being faster and more token-efficient.
SpaceXAI launches Grok 4.5 at $2/$6 per million tokens, targets coding and enterprise work
Elon Musk's SpaceXAI has released Grok 4.5, priced at $2 per million input tokens and $6 per million output tokens. The model, trained alongside recently-acquired Cursor, is positioned as a coding and enterprise tool that claims to outperform Claude Opus 4.8 on several benchmarks while undercutting it on price by 60-76%.
OpenAI releases GPT-5.6 with three variants after government security review
OpenAI is releasing GPT-5.6 to the public on July 9 following government security review under a Trump administration AI cybersecurity order. The release includes three variants: Sol (strongest), Terra (GPT-5.5 performance at half the cost), and Luna (lowest cost option).
Google Voice launches $10 and $20 monthly plans with call recording and Gemini transcription
Google Voice has introduced two paid subscription tiers available without a Google Workspace account. The Starter plan costs $10/month with call recording, while the $20/month Standard plan includes Gemini-powered call transcription and summarization.
Chinese AI Models Capture 30%+ of U.S. Developer Token Usage as OpenAI, Anthropic Costs Rise
Chinese AI models including DeepSeek and Z.ai have captured over 30% of weekly token usage by U.S. companies on OpenRouter since February 2025, up from 4.5% in the first half of the year. The shift comes as companies seek alternatives 60-90% cheaper than leading models from OpenAI and Anthropic, while Chinese models close the performance gap to within 6-9 months of U.S. frontier systems.
Google AI Plus at $4.99/month and AI Pro at $19.99/month expand Gemini context windows to 128K and 1M tokens
Google has detailed pricing and features for its Gemini app subscription tiers. AI Plus costs $4.99/month and includes 128,000 token context windows, while AI Pro at $19.99/month provides 1 million token context windows. Free users are limited to 32,000 tokens.
Claude Sonnet 5 ships with 1M token context and new tokenizer that increases costs 30-40% for English text
Anthropic released Claude Sonnet 5 with a 1 million token context window and 128,000 token maximum output. The model removes traditional sampling parameters and introduces a new tokenizer that generates approximately 30% more tokens than Sonnet 4.6 for the same English text—effectively a significant price increase despite unchanged nominal rates of $3/million input and $15/million output tokens.