Microsoft releases three multimodal AI models to compete with OpenAI and Google
Microsoft AI released three foundational models on April 2: MAI-Transcribe-1 for speech-to-text across 25 languages, MAI-Voice-1 for audio generation, and MAI-Image-2 for video generation. The company positions these models as cheaper alternatives to Google and OpenAI offerings. Models are available on Microsoft Foundry with pricing starting at $0.36 per hour for transcription.
Microsoft Releases Three Multimodal Models to Compete With OpenAI and Google
Microsoft AI announced the release of three foundational models on April 2, 2026: MAI-Transcribe-1, MAI-Voice-1, and MAI-Image-2. All three are now available on Microsoft Foundry, with the transcription and voice models also available in MAI Playground, a new large language model testing platform launched March 19.
The models were developed by Microsoft's MAI Superintelligence team, led by CEO Mustafa Suleyman. The team was formed and announced in November 2025.
Model Capabilities and Performance
MAI-Transcribe-1 converts speech to text across 25 languages and is 2.5 times faster than Microsoft's Azure Fast offering, according to the company. Pricing starts at $0.36 per hour.
MAI-Voice-1 generates audio, producing 60 seconds of audio output in one second. The model supports custom voice creation. Pricing begins at $22 per 1 million characters.
MAI-Image-2 is a video-generation model. Pricing starts at $5 per 1 million input tokens and $33 per 1 million output tokens.
Microsoft claims these models are cheaper than comparable offerings from Google and OpenAI, positioning cost as a primary competitive advantage in an increasingly crowded generative AI market.
Microsoft's Dual Strategy
The release reinforces Microsoft's strategy of building proprietary AI capabilities while maintaining its partnership with OpenAI. Microsoft has invested more than $13 billion in OpenAI through a multi-year agreement and integrates OpenAI models across its product portfolio.
According to Suleyman, a recent renegotiation of the Microsoft-OpenAI partnership enabled Microsoft to pursue independent superintelligence research. The company applies the same dual approach to semiconductors, both manufacturing its own chips and purchasing from external suppliers.
"We're building Humanist AI," Suleyman wrote in a blog post. "We have a distinct view when creating our AI models — putting humans at the center, optimizing for how people actually communicate, training for practical use."
Suleyman told VentureBeat that additional models from Microsoft AI will launch soon on Foundry and integrate directly into Microsoft products.
What This Means
Microsoft's move signals confidence in its ability to develop competitive foundation models independently while preserving strategic partnerships. The pricing structure—particularly the cheaper transcription and voice generation offerings—directly targets enterprises evaluating alternatives to established vendors. However, the company's continued reliance on OpenAI demonstrates that even with substantial internal AI capabilities, Microsoft views OpenAI's technology as complementary rather than redundant. The success of these models depends on adoption velocity and real-world performance matching the company's efficiency claims.
Related Articles
DeepSeek Releases Experimental V4-Flash-Vision-Exp, Claims Near-Parity With Opus 4.8 on Agent Benchmarks
DeepSeek has released V4-Flash-Vision-Exp, an experimental multimodal extension of V4-Flash that adds image understanding while preserving text reasoning capabilities. The company claims the model approaches or beats Anthropic's Opus 4.8 on its internal multimodal agent benchmarks.
DeepSeek Releases V4 Flash Vision Exp, an Experimental Multimodal MoE Model with 1M Context
DeepSeek has released V4 Flash Vision Exp, an experimental vision-enabled variant of DeepSeek V4 Flash 0731 that adds image understanding while matching the base model's text performance. The sparse mixture-of-experts model uses 13B active parameters out of 284B total and supports a 1M token context window.
Microsoft to Kill Excel's COPILOT() Function on September 14, 2026
Microsoft will shut down Excel's COPILOT() worksheet function on September 14, 2026, roughly a year after its preview launch. The company says the Copilot side pane already covers the same capabilities, so a planned 2027 general availability release has been scrapped.
Qwen Launches Qwen3.8 27B, an Open-Weight Vision-Language Model with 262K Context
Qwen has released Qwen3.8 27B, a 27-billion-parameter dense vision-language model with a 262K token context window, available now via OpenRouter at $0.45 per million input tokens and $3.20 per million output tokens.
Comments
Loading...