Microsoft on Friday shipped Microsoft Decision 1 through Foundry, entering the decision-model category that TypeSafe AI created three and a half weeks ago — and built it on Alibaba Group Holding's Qwen 3.5 9B rather than anything from its partner OpenAI1. The launch came three days after OpenAI opened its Decisions API to every developer in public beta.
Microsoft Chairman and CEO Satya Nadella announced the model on X. Achint Srivastava, vice president of software engineering in Microsoft's Office of the CTO, introduced Microsoft Decision 1 as a way to add decision-making to existing applications, agents, and workflows in a secure, trusted environment.
Base model and roadmap
Microsoft plans to rebase Microsoft Decision 1 on its own MAI models as well as OpenAI's. The model accepts up to 32,768 tokens of text and returns JSON but does not support images. Microsoft has yet to announce open weights. Cloudflare's Clef-flash uses the same Qwen 3.5 9B base Microsoft picked, while Cloudflare released CLEF under Apache-2.0 with a vision encoder.
Microsoft Decision 1 costs $0.042 per million input tokens with free output, matching what TypeSafe AI charges for Jev. Perplexity AI charges $0.02, and jaredpalmer lists Kev 4B on OpenRouter at $0.042 per million input tokens. OpenAI charges more than double Jev's rate.
Internal deployments and benchmarks
Microsoft said it is already testing Microsoft Decision 1 across the company. Xbox Research used it to sort more than 10,000 pieces of player feedback, and Microsoft said the model is more than 14 times as fast as GPT-6 Sol for that task. Microsoft Discovery used it to score experiments before an agent replans, with Microsoft claiming 46 times greater consistency. The Copilot team used it to grade chat and agent responses, and on-call engineers used it to pull context during live incidents. Microsoft lists model routing as a use case and has said GitHub Copilot will soon decide when to run a task on-device and when to send it to cloud-scale models, though Microsoft has not confirmed whether Microsoft Decision 1 will make routing calls for Copilot.
Microsoft says the model's probabilities are calibrated — a 90% prediction should be right about nine times in 10 on representative cases. Testing with eight kinds of perturbations, including reordered options and paraphrased descriptions, changed 1.3% of Microsoft Decision 1's answers on average. Three other scoring systems showed flip rates between 64.9% and 73.2% under the same approach.
TypeSafe AI's head start
TypeSafe AI announced an $870 million Series A at a $7.5 billion valuation led by Andreessen Horowitz on the same day Microsoft Decision 1 arrived. TypeSafe AI CEO Diogo Almeida posted on X that 29.4% of the Fortune 500 showed up in the three weeks since Jev's launch, though TypeSafe AI has not named any of them. Amazon Web Services, Upstage, and Ollama have adopted TypeSafe AI's System One API; Microsoft's Foundry sample code calls a /systemone endpoint, but Microsoft has not said whether Microsoft Decision 1 is fully compatible with that API.
A preprint called JevOut, led by USC computer science researcher Zixiang Xu, found that short, natural-sounding additions to context flipped 312 of Jev's 508 initially correct decisions, and in 229 cases Jev assigned at least 70% probability to the wrong answer. JevOut did not test Microsoft Decision 1.
ANALYSIS Microsoft's choice to ship on Qwen 3.5 9B rather than an OpenAI base underscores how quickly the decision-model category has moved: speed to market outweighed partnership alignment. The price-matching at $0.042 per million input tokens sets up a direct contest with TypeSafe AI for the same Azure enterprise customers that Diogo Almeida's Fortune 500 traction numbers represent.