Back Home

模型與開發者平台

Microsoft Launches MAI-Image-2.5-Pro and MAI-Voice-2-Flash, Using In-House Models to Separate the Quality and Cost Curves

Microsoft is bringing its highest-quality image generation model and a high-throughput voice model into public preview in Foundry simultaneously, targeting precision editing and real-time voice agents, respectively. The company has published cost and latency data from several production environments, but cross-provider comparisons still lack consistent, reproducible testing conditions.

Steven C. Price · CC BY-SA 4.0 · Image source
zh-Hant

Microsoft opened previews of MAI-Image-2.5-Pro and MAI-Voice-2-Flash on July 23, splitting the same model family into two distinct deployment profiles. Image-2.5-Pro targets hero images, localized editing, and text rendering in images. Foundry prices it at $5 per million text input tokens, $8 per million image input tokens, and $106 per million image output tokens. Voice-2-Flash is designed for customer service and real-time voice agents and costs $15 per million characters. Microsoft claims it is twice as fast and 32% cheaper than MAI-Voice-2 while retaining natural prosody and acoustic quality.

The technical significance of this update goes beyond adding new API model identifiers: Microsoft is beginning to use its in-house models across the full product data plane. Bing Image Creator now uses MAI-Image-2.5 exclusively by default. Following its rollout in OneDrive, Microsoft observed a 26% increase in save rate and an approximately 25% reduction in P95 latency. PowerPoint’s image workloads reportedly use up to 84% less GPU cost than GPT-Image-2, while Dynamics 365 Contact Center has achieved savings of up to 89% after switching to Voice-2-Flash. These figures come from real-world product telemetry and are therefore more relevant to engineering decisions than a single public leaderboard. However, Microsoft has not disclosed the traffic mix, hardware, quality thresholds, or model versions used for comparison.

The image model card shows that the base MAI-Image-2.5 model has an overall text-to-image Elo score of 1254 and a text-rendering score of 1278. The evaluation relies on blinded rater preferences and cannot directly establish the Pro version’s reliability for specific brands, Chinese-language fonts, or high-resolution editing workflows. Developers should next test time to first packet, concurrency limits, recovery from voice-stream interruptions, character consistency, and false positives from safety filters. They should also note that Pro’s image output price is substantially higher than that of the standard model in the same family.

Sources

  1. Introducing MAI-Image-2.5-Pro and MAI-Voice-2-Flash
  2. 微软最新模型 MAI-Image-2.5-Pro、MAI-Voice-2-Flash 发布