Microsoft AI models cost push targets OpenAI dependence
Microsoft put two in-house AI models in public preview and claimed lower GPU costs across Bing, PowerPoint, Dynamics and other products.
By Renata Fuchs · Policy Reporter
· 4 min read
Microsoft AI models cost less to run in several production workloads than third-party alternatives, the company claimed as it put two new in-house systems into public preview Wednesday. The releases, MAI-Image-2.5-Pro and MAI-Voice-2-Flash, are part of Microsoft’s push to route more product traffic through its own models rather than relying only on OpenAI’s frontier systems.
Microsoft AI’s Superintelligence team said its models are already used across Bing, PowerPoint, OneDrive, Dynamics 365, Excel, GitHub Copilot and Azure. The company did not disclose revenue tied to the new models or paid preview adoption, but it did publish unusually specific performance and cost claims for live products.
What are Microsoft's new AI models?
MAI-Image-2.5-Pro is Microsoft’s higher-fidelity image model for detailed generation, editing and in-image text. MAI-Voice-2-Flash is a speech model aimed at high-volume enterprise uses such as contact centers, voice agents and real-time speech applications.
Microsoft priced MAI-Image-2.5-Pro at $5 per million text input tokens, $8 per million image input tokens and $106 per million image output tokens. The company said the base MAI-Image-2.5 model recently ranked No. 2 for image editing on Arena, a community leaderboard for generative media systems.
MAI-Voice-2-Flash, first previewed at Microsoft Build, is priced at $15 per million characters. Microsoft said it runs twice as fast as MAI-Voice-2 and costs 32% less.
Where Microsoft says the savings show up
The strongest part of Microsoft’s case is not the model announcement itself, but the production data attached to it. Bing Image Creator now runs fully on MAI-Image-2.5, according to the company. In PowerPoint, Microsoft said MAI-Image-2.5 cuts GPU costs by as much as 84% compared with OpenAI’s GPT-Image-2.
In OneDrive, Microsoft said MAI-Image-2.5 is the default for key image-editing scenarios and has produced a 26% increase in save rates, about 25% lower P95 latency and 2.5 times greater efficiency under medium-utilization production workloads.
For voice, Microsoft said MAI-Voice-2-Flash now powers Dynamics 365 Contact Center, used by customers including T-Mobile and EasyJet, and can reduce GPU costs by up to 89%. The model is also available through Azure Voice Live for developers building speech-to-speech agents.
Microsoft also said Dragon Copilot, its healthcare product used by 170,000 medical providers and tied to 28 million patient encounters last quarter, now uses MAI-Transcribe-1.5 for multilingual workflows across 58 languages. The company said internal evaluations showed a 50% relative reduction in transcription and language-identification error rates across most languages.
How does this change Microsoft’s OpenAI position?
Microsoft CEO Satya Nadella framed the move on X as “Frontier Diffusion & Control,” saying Microsoft can use frontier models for frontier needs while serving high-volume product use cases with optimized internal models. He added that OpenAI and Anthropic models remain part of Microsoft’s orchestration system alongside MAI.
That framing matters because Reuters reported in April that Microsoft’s exclusive license to OpenAI’s technology had been revised into a non-exclusive arrangement. The Information reported last September that Microsoft had started using Anthropic models in some products. Microsoft is now presenting itself less as a single-model distributor and more as a router among first-party and partner models.
The company also described a “hill-climbing” method that trains smaller models around product-specific data, evaluations and workflows. Microsoft said MAI-Code-1-Flash delivered about a 10% higher code accept rate than GPT-5.4 Mini and Claude Haiku 4.5 in VS Code while using 10% fewer median tokens. It also said an Excel-tuned version reached parity with GPT-5.6 for common Excel tasks while running on Nvidia H100 and A100 GPUs.
Those numbers are Microsoft’s own measurements, not independent benchmarks. They are still directionally useful for the market: the company is arguing that the economics of AI features depend less on the largest model available and more on whether a task-specific model can meet quality thresholds at lower serving cost.
The new models are available in public preview through Microsoft Foundry and the MAI Playground. Microsoft said it is extending the same approach to Copilot Chat, Outlook and PowerPoint, which suggests the next phase is broader substitution of routine AI workloads with smaller internal models where the company believes quality is close enough.
This story draws on original reporting from VentureBeat.