Microsoft releases three multimodal AI models to compete with OpenAI and Google
Microsoft AI released three foundational models on April 2: MAI-Transcribe-1 for speech-to-text across 25 languages, MAI-Voice-1 for audio generation, and MAI-Image-2 for video generation. The company positions these models as cheaper alternatives to Google and OpenAI offerings. Models are available on Microsoft Foundry with pricing starting at $0.36 per hour for transcription.
Microsoft Releases Three Multimodal Models to Compete With OpenAI and Google
Microsoft AI announced the release of three foundational models on April 2, 2026: MAI-Transcribe-1, MAI-Voice-1, and MAI-Image-2. All three are now available on Microsoft Foundry, with the transcription and voice models also available in MAI Playground, a new large language model testing platform launched March 19.
The models were developed by Microsoft's MAI Superintelligence team, led by CEO Mustafa Suleyman. The team was formed and announced in November 2025.
Model Capabilities and Performance
MAI-Transcribe-1 converts speech to text across 25 languages and is 2.5 times faster than Microsoft's Azure Fast offering, according to the company. Pricing starts at $0.36 per hour.
MAI-Voice-1 generates audio, producing 60 seconds of audio output in one second. The model supports custom voice creation. Pricing begins at $22 per 1 million characters.
MAI-Image-2 is a video-generation model. Pricing starts at $5 per 1 million input tokens and $33 per 1 million output tokens.
Microsoft claims these models are cheaper than comparable offerings from Google and OpenAI, positioning cost as a primary competitive advantage in an increasingly crowded generative AI market.
Microsoft's Dual Strategy
The release reinforces Microsoft's strategy of building proprietary AI capabilities while maintaining its partnership with OpenAI. Microsoft has invested more than $13 billion in OpenAI through a multi-year agreement and integrates OpenAI models across its product portfolio.
According to Suleyman, a recent renegotiation of the Microsoft-OpenAI partnership enabled Microsoft to pursue independent superintelligence research. The company applies the same dual approach to semiconductors, both manufacturing its own chips and purchasing from external suppliers.
"We're building Humanist AI," Suleyman wrote in a blog post. "We have a distinct view when creating our AI models — putting humans at the center, optimizing for how people actually communicate, training for practical use."
Suleyman told VentureBeat that additional models from Microsoft AI will launch soon on Foundry and integrate directly into Microsoft products.
What This Means
Microsoft's move signals confidence in its ability to develop competitive foundation models independently while preserving strategic partnerships. The pricing structure—particularly the cheaper transcription and voice generation offerings—directly targets enterprises evaluating alternatives to established vendors. However, the company's continued reliance on OpenAI demonstrates that even with substantial internal AI capabilities, Microsoft views OpenAI's technology as complementary rather than redundant. The success of these models depends on adoption velocity and real-world performance matching the company's efficiency claims.
Related Articles
Microsoft's Copilot gets access to local Windows files and OS-level actions under 'Hybrid Intelligence'
Microsoft announced an upgrade to Copilot at its Windows and Surface event that gives the assistant access to local files and the ability to take actions across Windows. The company calls the underlying approach "Hybrid Intelligence," which combines local and cloud AI models. Pricing, model details, and availability were not disclosed in the available reporting.
Microsoft will turn Windows Search into a Copilot-connected command interface this fall
Microsoft announced a redesigned Windows Search menu that accepts short typed commands to change system settings and can converse with the new Copilot app without launching it. The company says the update arrives this fall. Model, pricing and availability details were not disclosed.
Liquid AI releases open d1-3B decision model: 16 ms on Jetson AGX Thor, 48.57 on Decision Index
Liquid AI released two open-weight decision models, d1-3B (text and image) and the experimental d1-omni-600M (text with image or audio). Unlike generative models, they answer in a single forward pass, and Liquid AI claims d1-3B scores 48.57 on its Decision Index 0.2.1, ahead of all 4B and 9B models it tested.
Google releases EmbeddingGemma 2: 740M-parameter multimodal embedding model under Apache 2.0
Google announced EmbeddingGemma 2, a 740M-parameter natively multimodal embedding model built on the Gemma 4 architecture and released under Apache 2.0. Google says the quantized model needs about 191MB of active RAM for text-only weights and about 567MB for the full multimodal model on a Pixel 11 Pro. Google also launched a Mac app, AI Edge Foresight, to demonstrate it.
Comments
Loading...