Microsoft-Decision-1 lands in Foundry at $0.042 per million input tokens
A post-trained Qwen3.5-9B model that returns calibrated probabilities for fixed options instead of generating text. Available now in Foundry and OpenRouter.

A post-trained Qwen3.5-9B model that returns calibrated probabilities for fixed options instead of generating text. Available now in Foundry and OpenRouter.

Microsoft released Decision-1 on 9 October 2026. It scores closed-set choices at $0.042/M input with free output. Hosted only, Decisions API only.
Microsoft shipped Decision-1 on 9 October 2026. It is a decision-scoring model, not a chat model. You give it a state and a closed set of options. It returns a calibrated probability for each option in one pass. No generated text.
It is available now in Microsoft Foundry as a Direct from Azure model and on OpenRouter as microsoft/microsoft-decision-1. Weights are not open. You call the Decisions API, not the chat completions endpoint. Chat SDKs will not work. The Foundry model card lists it as generally available. OpenRouter shows it served by Azure with 99%+ uptime in the first days.
Base model is Qwen3.5-9B, post-trained by Microsoft for single-pass scoring. Context window is 32,768 tokens. Input is text only. The model supports yes/no (noul), multiple-choice, rating, classification, and rubric-based grading of AI responses or proposed agent actions. It can also check groundedness against evidence you supply in the state.
Output tokens are free. Input is $0.042 per million tokens on both Foundry (US and EU Datazone) and OpenRouter. OpenAI’s Luna decisions endpoint is $0.10 per million input. Jev sits at the same $0.042. At that price a million-token classification job costs four cents. High-volume routing or guardrail checks stop being a cost center.
Microsoft says it led a 36-benchmark suite of nearly 150,000 questions held out from training, with 83.5% average accuracy. Comparisons they published: 81.9% for Quyet-1.0-Large, 79.4% for GPT-6 Luna Decisions, 77.2% for H2O-Lightning-4B. They measured p50 latency around 85 ms in their chart—faster than the other decision models they tested and roughly 35× faster than GPT-6 Sol on the same tasks. One internal note puts it at 2.5× quicker than H2O-Lightning-4B v1.1.
Robustness tests showed decision flips on only 1.3% of perturbations on average. Zero flips when option descriptions were paraphrased or when options were reversed or shuffled. Calibration is claimed high enough that a 90% score should be right about nine times out of ten on representative cases. All numbers come from Microsoft’s own tests. It does not appear on the public JevBench board.
You deploy it from the Foundry model catalog. Search for Microsoft-Decision-1, deploy, then call your resource endpoint at /providers/microsoft/v1/systemone. The request takes a state (string, object, or array) and a set of typed questions. Each question has a type and instructions. Choice questions include criteria that map labels to descriptions. Score questions take an ordered list of levels.
Example pattern from the docs: state is a support ticket, questions ask which team should handle it and how urgent it is. Response returns the selected label, the full probability distribution, and a confidence signal. On OpenRouter the same Decisions API is used. Their playground already shows an example that gates a destructive delete_rows tool call on an 80% safety probability. If the noul score is below the threshold you set, the code refuses the action and routes to a human.
Rate limits, maximum options per question, and exact request schema limits are not published as of 10 October 2026. Authentication on Foundry uses Entra ID or API key. OpenRouter uses its standard key.
Xbox Research ran more than 10,000 open-ended feedback items from surveys, Steam, and X through it to bucket them into fixed themes. Quality was competitive with GPT-6 Sol while running over 14 times faster and 200 times less expensive.
Copilot team quality scoring of chat and agentic responses: competitive with GPT-5.6 Luna and 100 times faster.
On-call engineers retrieving relevant knowledge from logs, tickets, and messages: better and faster than an LLM for the retrieval step.
Microsoft Discovery adaptive replanning: the decision model scored experiment results against a rubric, revised the plan, and repeated. It was 46 times more consistent than the LLM scorer at three times the speed, cutting overall replanning time by nearly four times.
This is a control layer. It does not replace a generator. You still need a larger model when you want text, code, or open-ended reasoning. Decision-1 is for the points in a workflow where the next action is one of a known set: route this ticket, continue this agent step or stop, label this example, accept or reject this proposed tool call.
Because output is free and latency is low, you can afford to ask multiple questions in one call. A single request can score urgency, department, and frustration level at the same time. Your application then branches on the probabilities.
Limitations are clear. Closed weights, hosted only. No chat, no generation, no summarization, no translation. Benchmarks are vendor-run. Microsoft plans to rebase later versions on MAI and OpenAI models, but the current API shape is intended to stay the same.
For an indie building agents the practical value is the cost and speed of the gate. Every extra LLM call adds latency and dollars. A four-cent-per-million-token score that decides whether to act, route, or ask a human keeps the expensive model for the steps that actually need it. Test the thresholds on your own distribution before you trust them. The published accuracy and calibration numbers are Microsoft’s.
Source: https://commandline.microsoft.com/microsoft-decision-1-model-foundry/
OpenRouter: https://openrouter.ai/microsoft/microsoft-decision-1
Foundry docs: https://learn.microsoft.com/en-us/azure/foundry/foundry-models/how-to/use-foundry-models-microsoft-decision
At $0.042 per million input tokens and free output, a 4,000-token support ticket classification costs $0.000168. Ten thousand such tickets in a day cost $1.68. The same volume on a $0.10 decision endpoint would be about $4. On a full GPT-6 Sol call that also generates an explanation the cost and latency both jump. Microsoft’s internal Xbox example claimed 200× lower expense than GPT-6 Sol for theme labeling. Even if your own ratio is lower, the gap is large enough that high-volume loops stop being painful.
OpenRouter lists the model with a 32,768-token context and shows early traffic from apps doing URL classification, candidate search, and evaluation harnesses. Foundry pricing is identical for US and EU Datazone. No provisioned throughput units are required for the pay-as-you-go path.
A typical call looks like this. State is the raw input: “The API returns 500 on every call. Customer says refunds have been failing for three days.” Questions ask three things at once: is it urgent (noul), which team (choice with billing/engineering/support criteria), and priority score on a three-level rubric. The response returns three answer objects. Each carries the chosen label and the full probability vector. Your code can require urgency above 0.75 and billing probability above 0.6 before auto-escalating. Anything lower goes to a queue for human review. Because output is free you can ask the extra questions without a cost penalty.
The endpoint is not a chat completions route. On Foundry it is /providers/microsoft/v1/systemone. On OpenRouter it is /api/alpha/decisions. Standard OpenAI client libraries will fail. You need the Decisions client or a raw POST.
As of 10 October 2026 Microsoft has not published rate limits, maximum number of questions per call, maximum options per choice question, or regional availability beyond the Datazone note. The model card lists 32,768-token context. Exact parameter count beyond the Qwen3.5-9B base is not disclosed. No quantized weights, no self-host option, no offline mode. Safety testing is mentioned (5,250 requests across 11 benchmarks covering harmful content, jailbreaks, and prompt injection) but the detailed refusal rates are not public.
Agent loops need many small decisions: is this tool call safe, does this step need a human, which model should handle the next sub-task, is this answer grounded. Running a full frontier model for each of those checks is slow and expensive. A purpose-built scorer that returns calibrated probabilities lets the application set explicit thresholds instead of parsing free text. Microsoft is not the first; Jev, Cloudflare Clef, and others already occupy the category. Decision-1 is their hosted entry at a price that matches the cheapest public options and with distribution through both Foundry and OpenRouter on day one.
Satya Nadella posted the announcement on 9 October. The Command Line blog post is dated the same day. The Foundry catalog and OpenRouter listing both show the model live. That is the shipped product. No waitlist, no research-only preview for this one.
Test it on a slice of your own logs or tickets before you wire the threshold into production. Vendor accuracy numbers are a starting point, not a guarantee. The API shape is stable enough to build against today.
Community
0 comments
React to this article
Trending now
Written by
Kirtesh Admute
Founder
Kirtesh Admute is the founder of IndieFounder, a platform for founders, builders, and people curious about technology. He writes about AI, startups, software, product building, and the lessons that come from building in public.
See an issue with this story?
Continue reading