Microsoft evaluates DeepSeek V3 for Copilot to cut agent costs, will offer cheaper tier within weeks
Microsoft is evaluating a self-hosted version of DeepSeek V3 to power Copilot Cowork as agent costs spiral. The company plans to launch a lower-cost tier within weeks while moving to usage-based pricing, charging enterprises for actual compute consumed rather than flat fees.
Microsoft evaluates DeepSeek V3 for Copilot to cut agent costs, will offer cheaper tier within weeks
Microsoft is exploring a self-hosted, fine-tuned version of DeepSeek V3 or another open-source model to power Copilot Cowork, the agentic assistant in its Microsoft 365 suite, according to statements to Axios. The company expects to offer a lower-cost model tier within weeks.
The move comes as Microsoft simultaneously shifts Copilot Cowork to usage-based pricing, charging companies for actual compute consumption rather than flat subscription fees.
Agent economics force the shift
Agentic AI tools like Copilot Cowork call language models repeatedly as they work through tasks, creating costs that scale rapidly with usage.
"We have users who do hundreds of tasks a week, which is great, they're way productive, but the consequence is the costs can go very high," said Charles Lamanna, Microsoft's executive vice president for Copilot, agents and platform.
Copilot Cowork currently runs on Anthropic and OpenAI models. Both providers have raised prices and moved away from unlimited usage plans. Microsoft previously metered GitHub Copilot for similar cost reasons.
DeepSeek offers inference cost advantage
DeepSeek V3, released in December 2024, offers significantly lower inference costs compared to frontier models while maintaining competitive performance. The model is open-source and popular with developers seeking cost-effective options.
Microsoft says any DeepSeek deployment would be optional for customers and fully hosted on Azure, maintaining data inside Microsoft's cloud infrastructure under its security and compliance controls. The company claims it has fine-tuned the model and added safeguards, including bias reduction measures.
Political complications
The timing presents political challenges. Washington has discussed banning DeepSeek, sanctioned Chinese AI firms, and recently forced Anthropic to restrict its top models for non-US users in a dispute that required Commerce Department intervention.
Multi-model strategy emerges
The evaluation signals Microsoft's broader shift toward a multi-model approach, reducing dependence on any single AI lab. This marks a strategic change from its exclusive, often tense relationship with OpenAI.
Microsoft emphasized this remains an evaluation, not a final decision. The company will confirm its model selection when the cheaper tier launches.
What this means
The economics of agentic AI are forcing even Microsoft to reconsider its infrastructure choices. When a company with deep pockets and close ties to both OpenAI and Anthropic considers Chinese open-source models for cost reasons, it reveals how unsustainable current agent pricing has become at scale. Microsoft's willingness to publicly name DeepSeek as a candidate, despite the political environment, underscores the financial pressure. The shift to usage-based pricing and cheaper model options will likely accelerate across the industry as agents move from demos to production workloads.
Related Articles
OpenAI Launches ChatGPT for Financial Services to Automate Wall Street Analyst Work
OpenAI launched ChatGPT for Financial Services, a tailored enterprise product built with design partners Morgan Stanley and Evercore that automates research, financial analysis, and pitchbook creation. The tool, powered by GPT-6 Astra, targets tasks traditionally performed by Wall Street's junior analysts and associates.
Nvidia and Palantir Deploy AI to Manage Supply Chains, Starting With Nvidia's Own 1.3-Million-Part Racks
Nvidia and Palantir are integrating Nvidia's open Nemotron models and cuOpt optimization engine into Palantir's Foundry platform to manage complex supply chains. The first deployment is Nvidia's own operation, where a single Vera Rubin server rack contains 1.3 million components.
AWS Details How Amazon Bedrock Prompt Caching Cuts Input Token Costs by Up to 90%
Amazon Bedrock's prompt caching feature can cut input token costs by up to 90% on cache hits by storing repeated context like documents, system prompts, and tool definitions. AWS outlines six implementation patterns and pricing details, including a 25% premium for cache writes and 90% discount on cache reads.
Apple Launches Revamped Siri Powered by Google's Gemini, Excludes EU and China at Launch
Apple has released a beta of its rebuilt Siri, now powered by Google's Gemini models, as part of iOS 27 and related 2027 software updates. The assistant reads screen content and personal context but won't launch in the EU or China due to regulatory concerns.
Comments
Loading...