Microsoft released Microsoft-Decision-1, a 9-billion-parameter model that returns a probability for each of a fixed set of answer options instead of generating text. Per Microsoft, it runs in Microsoft Foundry and through OpenRouter, with input billed at $0.042 per million tokens and output free.
Microsoft says it post-trained Qwen3.5-9B for single-pass scoring and plans to rebase the model on others, including its own MAI family and OpenAI's. Microsoft-Decision-1 takes yes/no, multiple-choice and rating questions, and can grade an AI response or an agent's proposed action against a rubric. The target jobs are routing, classification, prioritization and checks inside agent workflows, where a slow model call at every step adds up. Microsoft's own arithmetic: 100 extra milliseconds across 20 sequential decisions costs two seconds.
The Foundry documentation spells out the shape of a call to Microsoft-Decision-1. An application sends a state, as text or JSON, plus typed questions: a true/false test, a choice among labeled options, or a position on an ordered scale. What comes back is numbers, with no written rationale, so a developer can set a threshold and send low-confidence cases to a person or a larger model.
The performance claims are Microsoft's, measured against models it picked. It reports that Microsoft-Decision-1 had the highest accuracy across 36 benchmarks, about 150,000 questions kept blind from training, and the fastest speed of the field: 2.5 times quicker than the runner-up H2O-Lightning-4B and 35 times quicker than GPT-6 Sol. It also says the model changes its answer on 1.3% of rephrased or reshuffled inputs, with no flips when options are reordered.
Internal trials supply the more concrete figures. Xbox Research used Microsoft-Decision-1 to sort more than 10,000 pieces of player feedback into researcher-defined themes, and Microsoft says quality stayed competitive with GPT-6 Sol at over 14 times the speed and a fraction of the cost. The Copilot team reported parity with GPT5.6 Luna for scoring responses, about 100 times faster.
For teams building agents, the appeal of Microsoft-Decision-1 is economic: a check before every tool call or retry only makes sense when the check is nearly free. The trade-off is opacity, since the model cannot explain a score. Microsoft's documentation tells developers to validate a probability threshold on their own labeled examples before wiring it into an automated gate.













