Skip to content
Dashboard

AI SDK evaluation to decisions: What changes in your code?

Content Engineer

AI SDK decisions use experimental_decide and provider .decisionModel() factories to answer typed questions about shared state. Gateway's older evaluation names remain deprecated aliases.

Direct OpenAI integrations need a separate behavior review because the provider now uses the native Decisions API with GPT-6 Luna, including native probability distributions.

Copy link to headingWhich names should you update?

Use the decision names in new code and update existing imports and model factories together. The API remains experimental, so keep a lockfile and check compatibility when upgrading packages, including patch releases.

Code or configuration

Current name or action

experimental_evaluate import from ai

experimental_decide

gateway.evaluationModel(id)

gateway.decisionModel(id)

Custom provider model aliases

Define them under decisionModels

Registry model lookup

registry.decisionModel('provider:model')

Type for a decision model argument

Experimental_DecisionModel

Type for an explicitly annotated registry

Experimental_DecisionProviderRegistry

Mock model for unit tests

Experimental_DecisionMockModelV4 from ai/test

Gateway HTTP endpoint

Keep POST /v1/evaluate

Check each provider's exported factory before renaming its code.

For example, @ai-sdk/typesafe-ai 3.0.12 exposes evaluationModel; AI SDK 7.0.129 accepts that model through experimental_decide by adapting its legacy method. Keep the factory available in your installed provider version until you upgrade to one that exposes decisionModel.

The deprecated Gateway aliases preserve the request and response shapes. For a Jev call through Gateway, the naming update retains state, questions, and the answers object. It does not require changing your choice labels or replacing your question definitions.

Before:

import { gateway } from '@ai-sdk/gateway';
import { experimental_evaluate as evaluate } from 'ai';
const model = gateway.evaluationModel('typesafe-ai/jev');
// Pass model, state, and questions to evaluate().

After:

import { gateway } from '@ai-sdk/gateway';
import { experimental_decide as decide } from 'ai';
const model = gateway.decisionModel('typesafe-ai/jev');
// Pass model, state, and questions to decide().

Keep HTTP clients on the documented /v1/evaluate route. Changing terminology in application code does not imply a new HTTP URL.

Copy link to headingWhat stays the same in a decision request?

Each decision answers named questions about one shared state. Choice questions select an option key, Score questions return a position on an ordered rubric, and Boolean questions return the probability that a statement is true. An array in state provides shared context for those questions; unrelated inputs need separate calls.

For example, a meeting application could decide what follow-up a set of notes needs. Install ai and @ai-sdk/gateway, then configure AI_GATEWAY_API_KEY or Vercel OIDC in the server environment:

npm install ai @ai-sdk/gateway

import { gateway } from '@ai-sdk/gateway';
import { experimental_decide as decide } from 'ai';
const result = await decide({
model: gateway.decisionModel('typesafe-ai/jev'),
state: {
notes:
'We agreed to publish the accessibility audit next week. ' +
'Maya will prepare the findings. No date was selected for review.',
},
questions: {
nextStep: {
type: 'choice',
instructions: 'What follow-up is needed before the audit can be published?',
criteria: {
schedule_review: 'The owner is known, but the review date is missing.',
assign_owner: 'No person is responsible for preparing the findings.',
ready: 'Both the owner and review date are established.',
clarify: 'The notes do not establish enough context to choose.',
},
},
hasOwner: {
type: 'boolean',
instructions: 'Has someone committed to preparing the audit findings?',
},
},
});
const answer = result.answers.nextStep;
console.log(answer.choice, answer.probabilities?.[answer.choice]);
console.log(result.answers.hasOwner.probability);

The application receives a proposed next step, not a scheduled meeting. Creating a calendar event would be a separate operation using the missing review date.

Decision results still include usage and provider metadata. Keep a guard around optional Choice and Score distributions: some providers adapt language-model output and return a structured judgment without a distribution. Boolean probability always refers to the statement being true, so a value near zero can be a strong negative answer.

Copy link to headingWhy does direct OpenAI code need a behavior review?

The OpenAI decision provider calls OpenAI's native Decisions API with gpt-6-luna. Older evaluation examples used a Responses-based structured-output adapter, which could judge data with Sol or Astra and accept Responses reasoning options.

openai.evaluationModel is itself a deprecated openai.decisionModel alias, so retaining the old name after upgrading doesn’t preserve the Responses integration either.

Concern

Native OpenAI decision behavior

Model

Use openai.decisionModel('gpt-6-luna')

Choice and Score output

Includes native probability distributions

Score meaning

Probability-weighted position on a zero-based rubric

Boolean output

Native predicate probability of true

Question execution

Each question is evaluated independently

Provider options

Supports safetyIdentifier; other OpenAI options produce warnings

Native confidence

Available by question ID in providerMetadata.openai.confidence

For a direct native call, install @ai-sdk/openai, set OPENAI_API_KEY, and use the provider's model instance:

import { openai } from '@ai-sdk/openai';
import { experimental_decide as decide } from 'ai';
const { answers } = await decide({
model: openai.decisionModel('gpt-6-luna'),
state: 'Maya will prepare the accessibility audit findings.',
questions: {
hasOwner: {
type: 'boolean',
instructions: 'Has someone committed to preparing the audit findings?',
},
},
});
console.log(answers.hasOwner.probability);

Remove providerOptions.openai.reasoningEffort from this native decision call.

If your workflow needs Sol or Astra to judge evidence with a chosen reasoning effort, use Responses structured output with generateText, Output.object, and openai.responses(modelId). Define the response schema and review the code that reads result.output.

Recheck acceptance thresholds after changing the integration. Native probabilities and a language model's prompted estimates have different origins; neither the shared field name nor a matching numeric threshold establishes equivalent behavior on your examples.

Copy link to headingHow do aliases and registries fit the new API?

Use a custom provider to give a model an application-specific name, then add a provider namespace with a registry. Callers can select that alias without repeating the underlying provider and model ID:

import { openai } from '@ai-sdk/openai';
import {
createProviderRegistry,
customProvider,
experimental_decide as decide,
} from 'ai';
const registry = createProviderRegistry({
meeting: customProvider({
decisionModels: {
followup: openai.decisionModel('gpt-6-luna'),
},
fallbackProvider: openai,
}),
});
const result = await decide({
model: registry.decisionModel('meeting:followup'),
state: 'Maya will prepare the accessibility audit findings.',
questions: {
hasOwner: {
type: 'boolean',
instructions: 'Has someone committed to preparing the audit findings?',
},
},
});
console.log(result.answers.hasOwner.probability);

Here, followup resolves to the registered Luna instance. fallbackProvider only resolves model IDs that are absent from the alias map. It does not retry a failed call, replace a model that lacks a question type, or inspect answer confidence.

Keep the inferred registry type so TypeScript retains its experimental decisionModel method. Use Experimental_DecisionProviderRegistry If you need an explicit annotation, the stable provider interfaces do not include the experimental decision method.

String model IDs default to AI Gateway. Set globalThis.AI_SDK_DEFAULT_PROVIDER to use another provider with a decisionModel method. Registered model instances keep their provider's credentials and configuration, so the examples above call OpenAI directly.

Copy link to headingWhat does the rename leave to your application?

AI SDK Core validates the questions and answers, but it does not select a backup model. Gateway conditional decision fallbacks can rerun a successful decision when an answer meets a configured uncertainty condition. Those policies belong in providerOptions.gateway.models, separate from a registry's fallbackProvider.

When a Gateway condition triggers a fallback, the second model receives all original questions and state, and Gateway returns its complete result. Both stages contribute to cost and latency, and the returned model may supply different probability fields. Keep those differences in the code that decides whether to accept a result or send it for review.

During migration, test the application branches with the new mock model and then compare real judgments against your labeled examples. Include missing distributions, invalid answers, and failed requests. Keep completed answers that call for clarification on a separate application path from requests that never produced usable answers.

Copy link to headingFrequently asked questions

Copy link to headingWill existing Gateway evaluation calls stop working because of the rename?

No. The experimental_evaluate export and gateway.evaluationModel factory remain deprecated aliases with the same request and response shapes. Use the decision names for new code and retain a lockfile because the API is experimental.

Copy link to headingShould I change the Gateway endpoint to /v1/decide?

No. Gateway continues to accept decision requests at POST /v1/evaluate. The SDK terminology change does not require an HTTP endpoint migration.

Copy link to headingCan I rename openai.evaluationModel to openai.decisionModel for Sol or Astra?

No. The current native OpenAI decision integration supports GPT-6 Luna. Use Responses structured output when the task requires Sol or Astra and review the schema, reasoning options, and result-reading code.

Copy link to headingDoes fallbackProvider retry a decision with another model?

No. The fallbackProvider setting resolves model IDs missing from the custom provider's alias map. Runtime fallback behavior needs a separate mechanism, such as AI Gateway's model fallback configuration.

Copy link to headingDo all decision models return probability distributions?

No. Native providers such as Jev and OpenAI Decisions return Choice and Score distributions, while structured language-model adapters can omit them. Check for a distribution before using its values in an application rule.

More Decision models articles

Ready to deploy?