Skip to content
Dashboard

What are decision models? How they work and when to use them

Content Engineer

Decision models are AI models that evaluate supplied context against defined questions and return structured answers, often with probabilities. They suit tasks such as assigning a support request to a team or rating an incident's urgency. Your application defines the allowed outcomes and how to use the prediction.

Imagine a customer writes, "I was charged twice. Please return the extra payment." The application requires a billing category and a refund-request flag before it can receive a written reply. You can use a decision model for those judgments and a generative model to draft the response.

Copy link to headingHow does a decision model work?

You supply state, the evidence the model should consider, along with questions and their answer definitions. State could contain a customer message, a document excerpt, or a record of an agent's actions. Each question should ask for one judgment that you can check against that evidence.

The model returns values your application can use in a branch, a queue, or a review process. Your code can use the billing label to select a support team and its associated probability to decide whether to route automatically or request review.

Native decision models can score the allowed outcomes without generating an answer token by token. Their implementations differ. Laya uses an encoder with a decision head, while Typical reads shared state into a cache and scores question-specific options. The category describes a task and interface; it doesn't imply that every provider uses the same architecture.

Several independent questions can share the same evidence. For example, Jev evaluates questions separately against one state. If a later question needs an earlier answer, your application must pass that result into a subsequent step.

Copy link to headingWhat kinds of answers can you ask for?

Three common question types cover different judgments. Liquid d1 and Jev support these types, and AI Gateway exposes them as Choice, Score, and Boolean.

Question type

Example

How to interpret the answer

Choice

Which queue should receive this ticket: billing, delivery, technical support, or review?

One selected option, with a distribution over options when the model provides it

Score

How urgent is the issue: routine, time-sensitive, or blocking work?

Position on an ordered rubric, which can fall between levels

Boolean

Is the customer asking for a refund?

The probability of a yes answer, between 0 and 1

TypeSafe-compatible APIs call the yes/no type Noul, while OpenAI calls it a predicate. Values near zero mean the model favors no; they don't signal uncertainty. An urgency score has a different meaning again: it describes the assessed urgency, rather than confidence in that assessment.

Define categories so that they separate the outcomes you care about. Include a review or insufficient-evidence option when a forced choice would hide missing information. If multiple labels can apply, ask separate yes/no questions. The same ticket can describe both a login problem and a refund request.

Copy link to headingHow do decision models differ from structured outputs and classifiers?

An LLM can already return a category in valid JSON. Structured Outputs constrains generated responses to a schema, including allowed enum values. The reason to consider a native decision model is its focused prediction interface and probability information, not the assumption that every LLM answer requires scraping prose.

Approach

What it gives your application

When to consider it

Rules or a decision table

An outcome derived from conditions you specify

The relevant facts and policy can be expressed exactly

Classifier

Labels or scores for an input

You have a defined classification task and suitable training data or a zero-shot model

Native decision model

Typed judgments over supplied questions, often with answer distributions

You need repeated, bounded judgments that application code can act on

LLM with structured output

Generated fields constrained by a schema

The task needs richer reasoning, extraction, or free-form content alongside a decision

Caller-defined labels also exist in zero-shot classifiers such as GLiClass. Classification and probability estimation have a longer history than the current decision-model products. Evaluate the actual question types, output semantics, and task accuracy instead of treating the product label as a technical guarantee.

Similarly, reward models learn to score outputs according to preferences. They overlap with decision models in evaluation work, but a preference score doesn't automatically represent the probability that a business condition is true.

An SDK can hide some of these differences. AI SDK's decision interface supports native Jev and OpenAI decisions with Choice and Score distributions. Its Anthropic and Google adapters use structured-output LLMs, omit those distributions, and return prompted Boolean probability estimates. An identical function call therefore doesn't establish identical model behavior.

Copy link to headingWhich decision models and APIs are available?

The category includes hosted APIs and models whose weights you can run yourself. These examples illustrate different implementations and access methods.

Model or service

Access

Defining capability

Liquid d1

Liquid API; AI Gateway as liquid/d1

Choice, Score, and Noul judgments over shared state without text generation

Jev from TypeSafe AI

Hosted API; AI Gateway

Typed questions with probability distributions and separate confidence values for Choice and Score

Perplexity's Decisions API

Hosted API; open model weights

Choice, Score, and Noul predictions over text and images

OpenAI's Decisions API

Public beta API

Luna-based Predicate, Choice, and Score answers over text or images, with native probabilities

Telnyx Decision Models

Hosted beta API

Choice, Noul, and Score questions through a TypeSafe-compatible interface

Mercury Decide from Inception

Hosted through OpenRouter

Choice, Score, and yes/no answers through a Jev-compatible request schema

Laya from Convai Innovations

Open weights; AI Gateway as convaiinnovations/laya

An encoder-based family with typed decision outputs and specialized checkpoints

Typical from Oz Labs

Open weights

Choice and Score distributions with an explicit abstention signal, plus yes/no probabilities

OpenAI introduced Decisions API at DevDay on September 29, 2026. The API is now in public beta, using gpt-6-luna through a dedicated Decisions endpoint. Its Choice and Score answers include distributions over the supplied options or levels.

For open-weight models, inspect the checkpoint's training scope before choosing it. Laya's model card, for example, distinguishes its base models from a checkpoint fine-tuned for typed-decision tasks. Benchmark results apply to the tested checkpoint and shouldn't be attributed to the whole family.

Copy link to headingWhat can you do with Perplexity's Decisions API?

Multimodal decision models can evaluate visual evidence alongside text. Perplexity's Decisions API uses pplx-decider-v1-27b to answer Choice, Score, and Noul questions about supplied state. You can ask several questions about the same content in one request.

For example, a support workflow could classify a customer's screenshot alongside their message to distinguish a payment error from a sign-in problem. Define an outcome for unreadable or insufficient evidence, then use the returned probabilities to decide whether to suggest a queue or request review. Identifying a payment error doesn't establish whether a charge settled; that requires the transaction record.

Perplexity also publishes model weights under Apache 2.0 and example inference code. Running the model yourself gives you control over its deployment, with responsibility for serving it and testing it on your workload. The local inference interface differs from the hosted API, so evaluate the integration as well as the predictions when choosing between them.

Copy link to headingWhen should you use a decision model?

Start with a repeated judgment whose answer space you can define and whose mistakes you can measure. Useful candidates include:

  • Routing a form submission to sales, support, or a review queue based on the message's intent.

  • Assessing an incident against an urgency rubric so responders can prioritize investigation.

  • Checking an agent trace for a specific outcome, such as whether the agent asked the customer for missing information.

  • Choosing whether a request needs a reasoning model or can follow an existing workflow.

Keep exact checks in code. If a payment record already has a trusted refunded field, reading that field is more direct than asking a model to infer it. Distinguishing a promised refund from a completed one in a conversation requires interpretation, which is where a model can help.

Existing decision tables can remain responsible for business policy. You can combine a model's interpretation of a refund request with explicit checks of the purchase date and the user's permissions. The request classification alone doesn't authorize a refund.

Use a generative model when you need a written explanation, a new plan, or source code. For a task that requires several deductions, test whether decomposing it into bounded questions preserves the necessary context. Multiple questions in one call don't automatically form a reasoning chain.

Copy link to headingHow can you try a decision model with AI Gateway?

AI Gateway lets you call Jev, Laya, and Liquid d1 through the same decision API. These examples ask each model whether a customer wants a refund, using the same message and question so you can compare the predictions.

Use AI SDK 7 or later, with AI_GATEWAY_API_KEY set in your server environment. Follow the decision quickstart for installation and credentials. The decision API is experimental.

Copy link to headingJev

import { experimental_evaluate as evaluate } from 'ai';
const result = await evaluate({
model: 'typesafe-ai/jev',
state: 'I was charged twice for one order. Please return the extra payment.',
questions: {
requestsRefund: {
type: 'boolean',
instructions: 'Is the customer asking for money to be returned?',
},
},
});
console.log(result.answers.requestsRefund.probability);

Copy link to headingLaya

import { experimental_evaluate as evaluate } from 'ai';
const result = await evaluate({
model: 'convaiinnovations/laya',
state: 'I was charged twice for one order. Please return the extra payment.',
questions: {
requestsRefund: {
type: 'boolean',
instructions: 'Is the customer asking for money to be returned?',
},
},
providerOptions: { gateway: { only: ['boundless'] } },
});
console.log(result.answers.requestsRefund.probability);

Copy link to headingLiquid d1

import { experimental_evaluate as evaluate } from 'ai';
const result = await evaluate({
model: 'liquid/d1',
state: 'I was charged twice for one order. Please return the extra payment.',
questions: {
requestsRefund: {
type: 'boolean',
instructions: 'Is the customer asking for money to be returned?',
},
},
});
console.log(result.answers.requestsRefund.probability);

Each example prints the estimated probability that the customer requested a refund, between 0 and 1. This prediction doesn't establish refund eligibility or confirm that a payment has been returned. Compare the models on representative messages before connecting their predictions to actions.

These calls use the AI SDK's Boolean probability field. TypeSafe-compatible responses expose the yes/no probability as noul. Liquid's direct API also uses a different model ID, d1:free, so keep the model ID and response fields matched to your integration.

For deployment differences and a shared product-catalog test, compare Jev, Laya, and Liquid d1.

Copy link to headingHow should you evaluate probabilities before automating a decision?

To use probabilities in application logic, first check how well they correspond to outcomes on your data. With a well-calibrated binary classifier, roughly 80% of cases assigned a probability near 0.8 should belong to the positive class. Fields called confidence can have a different meaning. Jev's confidence value summarizes the answer distribution; it isn't interchangeable with the selected option's probability.

Build a labeled test set containing routine cases, ambiguous requests, and missing evidence. Compare the model with your current approach using the same inputs. For support routing, measure both incorrect assignments and the share of tickets sent for review. Even a threshold with few mistakes may not improve the workflow if it sends almost everything to a person. Include fallback requests and review work when assessing the total cost and time to resolve a case.

Choose thresholds using those results and the consequences of each error. Mistaking a refund request for a shipping question has different consequences from letting a model approve a payment. Keep an explicit review outcome even when its probability is high: certainty that a case needs review is a reason to send it there.

Treat timeouts and invalid responses as failed requests, with their own fallback. AI Gateway also supports conditional decision fallbacks in beta. Choice and Score answers can trigger escalation on confidence conditions; Boolean answers use a probability range. Triggered fallbacks run and bill a second decision.

After changing a model, question, or category definition, rerun the labeled examples. The decisions can change even when the schema stays the same.

Copy link to headingFrequently asked questions

Copy link to headingAre decision models a replacement for LLMs?

No. Decision models fit judgments with defined outcomes, while generative models can produce explanations and other open-ended content. An application can use a decision model to select a workflow and an LLM to carry out the writing within it.

Copy link to headingIs a decision model the same as a classifier?

The tasks overlap, especially when choosing labels for text. Decision-model APIs often combine classification with ordered scoring and yes/no questions over shared context. Some zero-shot classifiers also accept labels at request time, so that capability alone doesn't distinguish the categories.

Copy link to headingDoes a high probability guarantee a correct answer?

No. Missing evidence or inputs that differ from the training data can lead a model to favor the wrong outcome with high probability. Measure calibration and errors on representative examples before using a threshold to trigger actions.

Copy link to headingDo all decision models support images?

Input support depends on the model and endpoint. Perplexity's Decisions API accepts text and images, while Jev uses textual state. Check the selected provider's input contract before passing media.

Copy link to headingCan you run a decision model yourself?

Yes. Open-weight options include Perplexity's pplx-decider-v1-27b, Laya, and Typical. Check the specific checkpoint's license, hardware requirements, and task coverage, then evaluate it on your application's examples.

More Decision models articles

Ready to deploy?