Skip to content
Dashboard

TypeSafe AI Jev vs. GPT-6 Astra: When should you use each?

Content Engineer

Use TypeSafe AI Jev for focused decisions with defined answers and native probabilities. Use GPT-6 Astra when the task also requires generating content or working through a broader problem with tools. Both can classify text into predefined categories. Choose between them by testing the decision your application needs.

For a support application, assigning a ticket and investigating the customer's problem are different jobs. Your model choice can reflect that separation, even when both jobs begin with the same message.

Copy link to headingHow do Jev and GPT-6 Astra differ?

Jev evaluates supplied state against typed questions. Its System One interface returns choices, rubric scores, or yes-or-no probabilities. It accepts text-based state, including structured records, and doesn't generate standard replies, explanations of its reasoning, or media.

GPT-6 Astra supports text generation, image input, and function calling. It also supports Structured Outputs, so you can constrain a classification to a set of labels.

Requirement

TypeSafe AI Jev

GPT-6 Astra

Select a support queue

Ask a choice question with category descriptions

Request a category through Responses Structured Outputs

Return a distribution over categories

Native choice answers include a probability for each option

The Responses schema in this example returns a label without a category distribution

Draft a customer reply

Use a separate generative model

Generate the reply from supplied context

Inspect an attached screenshot

Extract relevant text before evaluating it

Supply the image through a supported image-input API

Investigate using application tools

Application code retrieves evidence and acts on the decision

Configure tools for the model to call during the investigation

For classification, both approaches can produce a value your application understands. The useful comparison is how you define the question, what information accompanies the answer, and whether the same workflow needs other model capabilities.

Copy link to headingCan Astra return structured decisions like Jev?

Yes. OpenAI's Structured Outputs lets you define a schema for Astra's response. For ticket routing, that schema can restrict a team field to your queue names. The schema constrains the response format; a valid category can still be the wrong assignment.

Jev's Choice question defines the available categories and describes when each applies. Its answer includes the selected category and a distribution over those options. That distribution gives application code information to use when deciding whether to accept the selection or request review.

If you already classify tickets with Astra, a typed Jev answer alone isn't a reason to migrate. Check whether Jev's evaluation contract and probability outputs help you implement a decision rule that your application needs.

Copy link to headingHow would you compare the same ticket in AI SDK?

Consider this message:

Our invoice export fails every time. Signing in works, but downloading the file produces an error.

The word "invoice" could suggest billing, but the reported problem is a failed product feature. Define technical issues separately from questions about charges, then use the same descriptions for both models.

Use AI SDK decisions for Jev and Responses structured output for Astra, passing the same ticket and category definitions to both. Install ai, @ai-sdk/typesafe-ai@3.0.12, @ai-sdk/openai, and zod, then configure TYPESAFE_AI_API_KEY and OPENAI_API_KEY in the server environment. This TypeSafe provider version exposes evaluationModel, which experimental_decide accepts through its legacy-model compatibility support:

import { experimental_decide as decide, generateText, Output } from 'ai';
import { typeSafeAi } from '@ai-sdk/typesafe-ai';
import { openai } from '@ai-sdk/openai';
import { z } from 'zod';
const state =
'Our invoice export fails every time. Signing in works, ' +
'but downloading the file produces an error.';
const questions = {
team: {
type: 'choice' as const,
instructions: 'Choose the team that can resolve the reported problem.',
criteria: {
billing: 'Questions about amounts charged or payment status',
technical: 'Product features that fail or produce errors',
review: 'Unclear requests or problems requiring multiple teams',
},
},
};
const jev = await decide({
model: typeSafeAi.evaluationModel('jev-latest'),
state,
questions,
});
const astra = await generateText({
model: openai.responses('gpt-6-astra'),
prompt: JSON.stringify({ state, question: questions.team }),
output: Output.object({
schema: z.object({ team: z.enum(['billing', 'technical', 'review']) }),
}),
providerOptions: {
openai: { reasoningEffort: 'low' },
},
});
console.log(jev.answers.team.choice);
console.log(astra.output.team);

These calls use the providers directly. The Jev and AI SDK guide covers the alternative Gateway setup and application branching.

The example sets Astra's Responses reasoning effort to low, one of its supported values. Record that setting with the classification results so later comparisons use the same configuration. The schema restricts astra.output.team to the three queue names; Jev returns its selection under jev.answers.team.choice.

Under the supplied criteria, technical is the intended label for this authored example. Run representative tickets to compare actual selections. One clear example cannot establish which model fits the rest of your queue.

Copy link to headingDo the returned probabilities mean the same thing?

The examples produce different result fields. Jev returns a native decision, while Astra's Responses structured output returns the team field defined in the schema. That Astra schema requests a category without a probability distribution. AI SDK's openai.decisionModel uses the native Decisions API with Luna, so it cannot replace the Responses call for Astra.

Jev provides native probabilities. For the choice question above, you can inspect jev.answers.team.probabilities. TypeSafe also returns a separate confidence statistic, which AI SDK places in provider metadata. That statistic and the selected category's probability are different quantities.

If your application requires a category probability before assigning a ticket automatically, this Astra example does not supply it. Handle that absence explicitly in the routing rule. Assigning a missing probability a value of 1 would turn missing evidence into apparent certainty.

Asking Astra for a confidence number in a custom schema is another possible design, but a schema-valid number doesn't establish calibration. TypeSafe likewise explains that calibration applies across groups of predictions, without guaranteeing an individual answer. Check either model's judgments against labeled examples from your workflow before selecting thresholds.

Copy link to headingWhen should you keep Astra in the workflow?

Keep Astra when the requested output extends beyond the category. Suppose the ticket includes a screenshot and asks you to diagnose why the export failed. An investigation might require inspecting the image, retrieving logs, and explaining a fix. OpenAI documents Astra for workflows involving reasoning and tool use.

For this version of the task, a queue assignment is only the first step. You still need a component that can investigate the failure and produce a response.

Keeping Astra can also be reasonable when it already produces an acceptable classification as part of a larger response. Splitting that decision into another call adds an integration to maintain. Make the split when you need a separately evaluated routing step or a probability-based acceptance rule, then check that the resulting workflow meets your requirements.

If the investigation depends on account records, configure access to those records. The model name doesn't give Astra access to your billing database, and Jev can only evaluate the evidence your application supplies.

Copy link to headingWhen should you use GPT-6.1 Sol or GPT-6 Luna instead?

GPT-6.1 Sol is an alternative to Astra for complex coding and professional work, with lower token rates. For the failed-export ticket, test whether Sol can interpret diagnostic evidence and explain a supported fix. Compare the resulting investigations with Astra's on difficult cases, including incomplete logs or conflicting records, before deciding which model should handle them.

GPT-6 Luna targets focused tasks at high volume. Consider it when the next step is narrower, such as extracting the failing operation from a ticket or drafting a brief reply from a confirmed diagnosis. Both Sol and Luna support image input, Structured Outputs, and tool calling, so needing those capabilities alone doesn't require Astra. Measure the quality and cost of the completed task, including any retries or additional review.

Sol and Luna can also perform the classification through the Responses example by using openai.responses('gpt-6.1-sol') or openai.responses('gpt-6-luna'). Its schema still returns a selected label without a distribution. Luna also supports OpenAI's native Decisions API, which returns category distributions through openai.decisionModel('gpt-6-luna'). Switching to that API requires different request and answer handling, and new threshold testing.

Copy link to headingHow can Jev and Astra work together?

For the export ticket, your application could ask Jev to select a queue, then pass an accepted technical case to an Astra-powered assistant. The assistant would receive the original report and use any authorized diagnostic tools needed to complete an investigation.

Keep the original message in that handoff. Passing only technical would discard the failing operation and the customer's report that signing in still works. The classification should help select the workflow while preserving the evidence needed to complete it.

For uncertain cases, define what review means. You might send the original ticket to a person or ask Astra to assess it independently. If Astra is the reviewer, give it the category definitions and original evidence so it can make its own assessment.

Then evaluate the combined workflow. Even with a correct Jev classification, an unsupported Astra explanation leaves the customer with a bad answer. Assess routing and resolution separately so you can locate the failure.

Copy link to headingWhat should you test before switching models?

Use a labeled ticket set with the same category definitions for both implementations.

Include cases that challenge the distinction between billing and technical work:

Test case

Intended behavior

What to inspect

Invoice export fails with an error

Select technical

Whether the model follows the reported failure rather than the invoice keyword

Export works, but the charged amount is disputed

Select billing

Whether the label reflects the payment question

Export fails and the customer disputes a charge

Select review under the example's criteria

Whether the workflow preserves both requests

Message says only "invoice problem"

Select review

Whether missing context leads to a request for clarification

Classification request fails

Keep the ticket available for review or retry

Whether the application avoids dropping or silently assigning the ticket

Record the model identifier and request configuration with each result. Preserve the question wording, including any revisions to the category descriptions. For Astra, include reasoning effort so later comparisons use the same setup.

Compare routing mistakes with the amount of work sent for review. If you introduce an acceptance threshold for Jev, inspect the cases it rejects as well as those it accepts. For a workflow that also drafts replies, have reviewers assess whether the final answer addresses the original request and uses supported facts.

Copy link to headingWhat should neither model decide on its own?

Neither a queue label nor a generated recommendation should authorize a refund. Your application must check the relevant account facts and permissions before executing an action.

Keep classification separate from claims about what happened. Asking for a refund doesn't establish that a refund was issued. Read the transaction record when the workflow needs that fact, and handle unavailable evidence explicitly.

Copy link to headingFrequently asked questions

Copy link to headingCan GPT-6 Astra classify tickets into a fixed set of categories?

Yes. Astra supports Structured Outputs, which can restrict an answer to category names in a schema. In the AI SDK example, generateText and Output.object return the allowed team value through the Responses API.

Copy link to headingDoes Jev replace GPT-6 Astra in a support assistant?

Jev can take responsibility for a defined decision such as selecting a queue. If the support assistant also investigates issues or writes replies, Astra can continue handling that work.

Copy link to headingCan I reuse a Jev probability threshold with Astra?

No. The Astra schema in this example returns a team label without a category distribution. Adding a prompted probability estimate to the schema would still require testing a new acceptance rule against labeled tickets.

Copy link to headingWhy does the Astra example specify low reasoning effort?

The example explicitly selects a supported Responses reasoning effort so its configuration can be reproduced. Test low against other supported settings on your tickets before choosing an effort for production.

Copy link to headingDoes structured output guarantee the right routing decision?

No. Restricting the answer to an allowed category controls its format. Both Jev and Astra still need evaluation against examples with known destinations to assess whether their judgments fit your routing policy.

Copy link to headingCan I use GPT-6.1 Sol or GPT-6 Luna in the AI SDK example?

Yes. In the Responses classification call, replace openai.responses('gpt-6-astra') with openai.responses('gpt-6.1-sol') or openai.responses('gpt-6-luna'). Both support the example's explicit low reasoning effort. Keep the same ticket, schema, and category definitions when comparing their routing decisions.

Copy link to headingDo GPT-6.1 Sol and GPT-6 Luna support the same reasoning settings?

Both support the Responses reasoning settings low, medium, high, xhigh, and max, but only Luna also supports none. For a comparison that isolates the model choice, use an effort supported by both before tuning each model separately. Native Luna Decisions does not accept Responses reasoning options.

Copy link to headingNext steps

More Decision models articles

Ready to deploy?