OpenAI Decisions API and TypeSafe AI's Jev both return choices, rubric scores, and probabilities for yes-or-no questions. Decisions API is in public beta and accepts text and images; Jev accepts text-based state. For an existing Jev application, compare the input format, answer fields, and thresholds its code uses.
Two models can select the same category without providing the same information for deciding whether to act on it. Before replacing a decision service, inspect the data your application reads after the answer arrives.
Copy link to headingWhich capabilities do both products provide?
Both APIs provide typed answers and native probability information. Their question names and request formats differ, so a migration needs to preserve what each question means as well as adapt its syntax.
Sources: OpenAI Decisions guide, Jev input formats, and TypeSafe's question types.
For direct HTTP integrations, OpenAI takes shared input and a questions array, then returns an answers array with question names. Jev uses state and maps questions and answers by their IDs. OpenAI choice probabilities are an array of value-probability pairs; Jev choice probabilities are an object keyed by option. Adapt the result-reading code before comparing model behavior.
Copy link to headingWhy do probabilities and confidence matter for a Jev migration?
Each Jev Choice answer includes the selected option, probabilities for the available options, and a confidence value. OpenAI choice answers provide those kinds of information too. The selected option might identify a support queue, while the distribution helps you detect a request that spans two teams.
TypeSafe's confidence value summarizes the shape of that distribution. It differs from the probability assigned to the selected option. Probability spread across several answers produces lower confidence than probability concentrated on one answer. OpenAI also returns a separate confidence field, but matching field names do not establish equivalent calculations or calibration.
Suppose your application sends a ticket for review when Billing and Account access both receive substantial probability. The OpenAI response provides the option probabilities needed to express that rule. You still need to test whether the existing thresholds send the right requests for review with the new model.
Do not carry a numeric threshold across providers because both outputs fall between zero and one. Test the acceptance rule again against labeled examples, including the cases it sends for review. The threshold is useful only when the resulting routing behavior meets your requirements.
Jev's other question types also need explicit mappings. Score and OpenAI score both evaluate ordered levels, so preserve the rubric descriptions and their order. Noul maps conceptually to OpenAI predicate: both estimate whether a condition is true. Values near zero express a negative answer, not low confidence in a positive one.
Copy link to headingWhen does image input change the comparison?
Decisions API can evaluate visual evidence alongside text. Jev's state format is text-only, so a Jev workflow needs relevant visual information converted into text before evaluation.
Consider a user who reports that an export failed and attaches a screenshot. The message may omit the error visible in the image. With Decisions API, the application can include that screenshot as an inline base64 data URL when choosing the next support step. Hosted image URLs and file_id inputs are unsupported. With Jev, the application first extracts the error text and supplies it with the message.
Compare those complete workflows using the same original inputs. Check whether text extraction loses information needed for the decision, and include that extraction step when measuring elapsed time or reviewing failures. Direct image input is a useful capability to test; it doesn't establish that one workflow will make better assignments.
If your application already has the error as structured text, use that record in the comparison. There is no need to add a screenshot-processing step to a decision whose evidence is available in the application.
Copy link to headingDoes AI SDK make the two interfaces interchangeable?
AI SDK decisions gives both providers a shared interface through experimental_decide. Use openai.decisionModel('gpt-6-luna') for native OpenAI decisions. With @ai-sdk/typesafe-ai 3.0.12, use the evaluationModel factory with the model ID jev-latest; experimental_decide supports that legacy interface. Both integrations supply native Choice, Score, and Boolean answers, including probability distributions for Choice and Score.
The SDK accepts shared state and questions keyed by ID, then returns answers under those IDs. Its boolean question maps to the providers' yes-or-no probability answers. This reduces the request and response changes needed when comparing providers, but it doesn't make their judgments or thresholds equivalent.
Provider-specific confidence statistics are separate from the SDK's answer probabilities. For Jev, the typesafe entry in result.providerMetadata contains confidence values keyed by question ID. Code that depends on this metadata needs a provider-aware rule; do not assume another provider supplies the same statistic there.
The current decision API is experimental and can change in patch releases. Older AI SDK evaluation examples need more than a function rename: OpenAI decision models now call the native endpoint and do not accept Responses API reasoning options. The SDK's documented shared state supports text and JSON values, so use the direct OpenAI input contract when you need image decisions.
For a workflow that needs Sol or Astra to generate a schema-defined review, the Sol with AI SDK guide uses Responses structured output.
Copy link to headingWhat would justify changing an existing Jev workflow?
Changing providers should improve the task your application performs. Start with one decision, such as assigning the first support team, and preserve its category definitions. Run the same reviewed cases through each candidate with the same inputs.
Inspect errors by category. Confusing Billing with Technical support may create extra handling time; misclassifying a request that needs review may send an unresolved problem into an automated path. An overall accuracy score can hide that difference. Record how often each workflow defers a case and whether the remaining automated assignments are acceptable.
Keep the downstream action separate during evaluation. You can compare proposed queue assignments without moving live tickets. After selecting a provider, test failures and timeouts as well as successful answers so the application preserves requests when a decision call cannot finish. Direct OpenAI calls can return a refusal for an individual question, so check each answer type before using its fields.
Neither a valid label nor an uncertainty value guarantees the underlying judgment. If the task expands into investigating the customer's problem and writing a response, evaluate that work separately. Vercel's Jev and GPT-6 Astra comparison covers where a general-purpose model fits alongside a decision step.
Copy link to headingFrequently asked questions
Copy link to headingIs Decisions API a drop-in replacement for Jev?
No. Both APIs return choices, scores, and probabilities, but their direct request and response formats differ. AI SDK provides a shared decision interface; existing acceptance thresholds still need testing against the replacement model.
Copy link to headingIs Jev's confidence the probability that its chosen answer is correct?
TypeSafe defines confidence as a summary of the distribution across answer options. The selected option also has its own probability, so code should distinguish the two fields and validate any acceptance rule against reviewed cases.
Copy link to headingCan both products evaluate screenshots directly?
Decisions API accepts inline base64 image data URLs alongside text. Jev accepts text-based state, so a screenshot workflow needs to extract the relevant information before sending it to Jev.
Copy link to headingDo decision APIs remove the need for a general-purpose model?
No. Defined choices can handle one step in an application, such as selecting a support queue. Work that also requires a written explanation or further investigation still needs a component capable of that work.