THE SHORT VERSION
● $0.042 per million input tokens, output free. OpenAI's Decisions API is $0.10. Same structure, 58% less.
● Built on Alibaba's Qwen3.5-9B, post-trained for decision scoring. Microsoft's own partner shipped the competing product three days earlier.
● It publishes accuracy AND calibration — the numbers nobody had last week. On Microsoft's own benchmark, naturally.
● 32,768-token context, text only, proprietary. No open weights despite the open base model.
What it is
Decision-1 scores a closed set of options and returns a calibrated probability for each, as JSON, from a single call. Yes/no, multiple choice, rating, classification, rubric questions, grading another model's response, groundedness checks against supplied evidence, and explicit abstention.
It is available on Microsoft Foundry as a "Direct from Azure" model, and on OpenRouter — where, notably, it runs on the Decisions API rather than the chat endpoint.
The limitations in the model card are unusually direct: not for text generation, chat, translation or summarisation; it produces no explanations or rationales; and it is not to be the sole automated decision on credit, employment, housing, healthcare or legal rights.
Four products, six days
| Product | Input / 1M | Output | Shipped |
| OpenAI Decisions API | $0.10 | Free | 6 Oct |
| Liquid AI d1-3B | Open weights | Zero tokens | 8 Oct |
| TypeSafe Jev | $0.042 | Free | Early access |
| Microsoft Decision-1 | $0.042 | Free | 9 Oct |
Microsoft priced to the cent against Jev, not against OpenAI. That tells you who it thinks the competition is — and it is not the company it owns a large stake in.
The numbers nobody had
Last week the honest answer to "which structured-decision model is more accurate" was that nobody had published a comparison. Microsoft has now published one, across 36 benchmarks and 147,137 questions:
| Model | Accuracy | Calibration |
| Microsoft-Decision-1 | 83.5% | 92.2 |
| Quyet-1.0-Large | 81.9% | 93.1 |
| GPT-6 Luna Decisions | 79.4% | 89.9 |
| H2O-Lightning-4B | 77.2% | 91.8 |
Two things to hold at once. This is the first published head-to-head on the metric that actually matters for these products, and it is worth having. And it is Microsoft's benchmark, on which Microsoft wins — which is what vendor benchmarks are for.
The detail worth noticing is the one Microsoft did not win: Quyet-1.0-Large beats it on calibration, 93.1 to 92.2. For a product whose entire proposition is that a 0.7 means 0.7, calibration is arguably the more important column, and Microsoft published a table where it comes second in it. That is a small point in the table's favour.
The speed claim, and why to discount it
Microsoft reports p50 latency of 85ms and p95 of 125ms through Foundry, and says that is 4.5x faster than Quyet-1.0-Large and 35x faster than GPT-6 Sol at 3.01 seconds.
That comparison is contested. Competitor figures use JevBench's adjusted median, and H2O.ai argues the method inflates measured times. Decision-1 is not on the JevBench board at all.
So: the absolute latency figures are Microsoft measuring its own model on its own infrastructure, and the multiples against competitors rest on a methodology one competitor disputes. Useful as a signal that it is fast. Not useful as a number.
The strategic part
Microsoft is OpenAI's largest partner. OpenAI launched the Decisions API on 6 October. Three days later Microsoft shipped a competing product, built on a Chinese open-weight model, at 58% of the price, and published a benchmark showing it more accurate than OpenAI's.
None of that requires a dramatic reading. Microsoft sells Azure, Azure sells inference, and a cheap high-volume classification endpoint is good Azure business whoever built the weights. But it is a clear statement that the partnership does not extend to leaving a market segment alone.
The more interesting signal is the base model. A Qwen derivative shipping as a first-party Microsoft product, generally available on Foundry, is a larger shift in what counts as an enterprise-grade model than the pricing is.
What to do
- Already on OpenAI's Decisions API? Decision-1 is 58% cheaper with the same billing shape. A migration test is a day of work on your own labelled data.
- Need images? OpenAI, still. Decision-1 is text only, as is Jev.
- Need long inputs? OpenAI, still. 32,768 tokens is a real constraint against a 1M-token context.
- Need an explanation with the answer? None of them. The model card says so outright.
- Regulated decisions? Microsoft explicitly excludes credit, employment, housing, healthcare and legal rights as sole automated decisions. Read that before you design around it.
Full three-way comparison with pricing and limits: Decisions API vs Jev vs Decision-1.
Sources