Skip to content
Dashboard

Using GPT-6.1 Sol to evaluate application data with AI SDK

Content Engineer

GPT-6.1 Sol can evaluate application data through AI SDK's structured-output support. Supply the evidence and a response schema to classify a proposed reply, score it against a rubric, or estimate whether a statement is true. Use the Responses API for Sol and Astra; OpenAI's native Decisions integration supports Luna.

Copy link to headingWhich AI SDK API should you use to evaluate data with Sol?

Use generateText with Output.object to ask Sol for a schema-validated review. For example, a support application can assess a draft reply against the customer's request and the relevant policy before showing it to a reviewer.

The evaluation happens during the application's workflow. Measuring how accurately Sol performs that evaluation requires a separate comparison with examples that people have already reviewed.

The examples below define three review fields in an application schema:

Review field

What you define

What the generated answer contains

Disposition

Named options and their meanings

One option key in disposition

Coverage

An ordered rubric with three levels

Fractional position in coverage, from zero to two

Unsupported claim

Statement to assess, with true and false criteria

An estimated probability in unsupportedClaimProbability

These fields are application-defined structured output. They do not include native probability distributions, and a generated estimate of 0.9 is not proof that the model is correct nine times out of ten on your task.

OpenAI's Decisions API provides native decisions with GPT-6 Luna. AI SDK's openai.decisionModel now calls that endpoint, and openai.evaluationModel is a deprecated alias of the same factory. To keep using Sol or Astra for these reviews, the examples below call Responses directly with generateText and Output.object.

Copy link to headingWhen is Sol useful for an evaluation step?

GPT-6.1 Sol supports complex coding and professional workflows. It is worth testing when a decision requires interpreting several pieces of evidence, such as deciding whether a draft explanation accurately describes a code change or follows a policy with exceptions.

Consider a customer asking whether an unused purchase qualifies for a return. The proposed reply says the refund has already been processed. Even if the customer is eligible, the policy alone cannot support a claim that money has been returned. An evaluation can separate whether the reply addresses the request from whether its factual claims are supported.

Supply the customer message, the applicable policy, and the draft reply together. If the application has a payment record, include it as a distinct source. Keeping those inputs separate helps define what the model should judge and gives a reviewer a way to investigate a disputed result.

Copy link to headingHow do you evaluate a draft reply with Sol?

Copy link to heading1. Install the AI SDK and OpenAI provider

In a TypeScript application, install the AI SDK, OpenAI provider, and Zod:

npm install ai @ai-sdk/openai zod

Configure OPENAI_API_KEY as an environment variable available to the backend code that makes the request. These examples use openai.responses() directly, so the model ID is gpt-6.1-sol without a Gateway prefix.

Keep a lockfile to make the package versions reproducible, and check the structured-output contract when updating the SDK.

Copy link to heading2. Define the evidence, criteria, and output schema

Put the shared example in src/reply-review.ts. The policy below is sample application data, and the questions describe how to assess a reply against it.

src/reply-review.ts
import { z } from 'zod';
export const reviewState = {
policy: 'Unused purchases can be returned within 30 days of delivery.',
customerMessage:
'My order arrived 12 days ago and is unopened. Can I return it?',
draftReply:
'Your purchase qualifies for a return. We have processed your refund.',
};
export const reviewQuestions = {
disposition: {
type: 'choice' as const,
instructions: 'Assess the draft reply using only the supplied evidence.',
criteria: {
ready: 'The reply answers the request and all factual claims are supported.',
revise:
'The reply contains an unsupported claim, contradicts the evidence, or omits part of the request.',
needs_review:
'The request or applicable policy is too ambiguous to assess whether the reply answers it.',
},
},
coverage: {
type: 'score' as const,
instructions: 'How fully does the reply address the customer request?',
criteria: [
'Does not address the request',
'Addresses part of the request',
'Addresses the full request',
],
},
unsupportedClaim: {
type: 'boolean' as const,
instructions:
'Does the reply make a factual claim not supported by the supplied evidence?',
criteria: {
true: 'At least one claim lacks support or contradicts the evidence.',
false: 'Every factual claim is supported by the evidence.',
},
},
};
export const reviewSchema = z.object({
disposition: z.enum(['ready', 'revise', 'needs_review']),
coverage: z.number().min(0).max(2),
unsupportedClaimProbability: z.number().min(0).max(1),
});
export const reviewPrompt = JSON.stringify({
task: 'Assess the draft reply using the evidence and question criteria.',
state: reviewState,
questions: reviewQuestions,
outputInstructions: {
disposition: 'Return the selected disposition key.',
coverage: 'Return a fractional rubric position from 0 through 2.',
unsupportedClaimProbability:
'Estimate the probability that unsupportedClaim is true, from 0 through 1.',
},
});

disposition provides a recommendation for the review workflow, while coverage measures how much of the request the reply addresses. unsupportedClaim focuses on factual support. Coverage alone should not determine whether a reply is ready because it may address the whole request and still make an unsupported promise.

The Responses call evaluates these criteria together in one prompt. If you change a question or its criteria, rerun the whole set against your reviewed examples to check whether the judgments remain useful.

Copy link to heading3. Call Sol and read the structured review

Import the shared inputs from a module that runs in your application's backend:

src/evaluate-sol.ts
import { openai } from '@ai-sdk/openai';
import { generateText, Output } from 'ai';
import { reviewPrompt, reviewSchema } from './reply-review';
const result = await generateText({
model: openai.responses('gpt-6.1-sol'),
prompt: reviewPrompt,
output: Output.object({ schema: reviewSchema }),
providerOptions: {
openai: { reasoningEffort: 'medium' },
},
});
console.log(result.output.disposition);
console.log(result.output.coverage);
console.log(result.output.unsupportedClaimProbability);

The validated fields appear under output, with disposition containing one of the three configured option keys. The coverage score falls between 0 and 2 for this rubric and may be fractional. unsupportedClaimProbability is the model's prompted estimate for that specific statement, rather than confidence in the entire review.

GPT-6.1 Sol supports reasoning effort settings of low, medium, high, xhigh, and max. It does not support none or minimal. This Responses example explicitly selects medium. Compare effort settings on cases that require different amounts of interpretation before choosing a default for your workload.

Copy link to headingHow should your application use the result?

Display the disposition beside the original evidence and draft in the review interface. Route revise results back to the author for correction, and use needs_review to request an additional record or a person's judgment. The ready option means the model found the reply acceptable under the supplied criteria. Sending it to the customer remains a separate operation.

Keep the supporting fields visible when they disagree. If the disposition is ready but the unsupported-claim estimate is high, the reviewer needs to see that conflict. Decide how your application handles such combinations using labeled examples, and store the question definitions with the evaluation so you can reproduce what was asked.

Avoid converting the coverage score into a probability. Its scale describes positions on your rubric, and changing the rubric changes the score's meaning. If you add a fourth level, update both the schema and prompt so the upper bound becomes 3.

Copy link to headingCan I use Astra or Luna for the same review?

Yes. Use the same Responses structured-output approach with a compatible model. GPT-6 Astra and GPT-6 Luna are alternatives you can compare using the same evidence and questions.

Copy link to headingUse Astra for complex reasoning across supplied evidence

Astra is a candidate when the review involves difficult reasoning across several records. Include those records in the shared state before making the call, since this evaluation assesses the supplied evidence without retrieving additional information.

src/evaluate-astra.ts
import { openai } from '@ai-sdk/openai';
import { generateText, Output } from 'ai';
import { reviewPrompt, reviewSchema } from './reply-review';
const result = await generateText({
model: openai.responses('gpt-6-astra'),
prompt: reviewPrompt,
output: Output.object({ schema: reviewSchema }),
providerOptions: {
openai: { reasoningEffort: 'high' },
},
});
console.log(result.output.disposition);

Like GPT-6.1 Sol, Astra supports reasoning efforts from low through max and does not accept none or minimal. The example chooses high for a demanding review; test whether that effort improves the judgments your application needs.

Copy link to headingUse Luna for focused evaluations repeated across many inputs

Luna is a candidate for repeated checks with a narrow rubric, such as determining whether replies address a clearly stated request. Include ambiguous and incomplete examples when testing it, since routine successes can conceal weaknesses on the cases that need review.

src/evaluate-luna.ts
import { openai } from '@ai-sdk/openai';
import { generateText, Output } from 'ai';
import { reviewPrompt, reviewSchema } from './reply-review';
const result = await generateText({
model: openai.responses('gpt-6-luna'),
prompt: reviewPrompt,
output: Output.object({ schema: reviewSchema }),
providerOptions: {
openai: { reasoningEffort: 'none' },
},
});
console.log(result.output.disposition);

Luna supports none on the Responses API, and you can raise the effort if testing shows that reasoning improves the judgments you need. These examples demonstrate different configurations. For a comparison that isolates the model choice, use the same supported effort setting across all three, then tune each model separately.

Copy link to headingWhat happens if an evaluation cannot return usable answers?

Handle request failures and failures to produce schema-valid output. generateText rejects output it cannot parse or validate, and accessing result.output can throw if no output was produced. Your application should record that failure separately from a completed evaluation whose disposition is needs_review. The latter is a judgment about the supplied evidence; the former means you have no usable evaluation result.

Use abortSignal when a request needs a deadline or cancellation, and configure maxRetries for transient failures. Retries can recover from temporary provider errors; missing evidence needs to be added to reviewState and included in a new prompt before another evaluation can assess it. Preserve the draft and offer another review path if the call fails.

The examples use a complete, non-streaming result. Keep one reply and its evidence in each review prompt so every result has an unambiguous subject.

Copy link to headingHow do you check whether the evaluations are useful?

Build a set of draft replies with expected dispositions established by reviewers. Include cases where the reply is fluent but unsupported, as well as cases where the policy itself is incomplete. Track false approvals separately from unnecessary revision requests because those mistakes have different consequences for the support team.

Compare the model's judgments with those labels after changing the model, reasoning effort, or question definitions. Record request duration and token usage alongside the decisions to assess the cost of running the check across your expected volume. Keep a separate test for how application code responds to each answer and to failures; correct branching does not prove that the model chose the right answer.

If your application needs native distributions over a defined set of answers, Jev with AI SDK provides another decision path. OpenAI's native Luna Decisions integration also returns distributions. Native decision probabilities and the generated estimate in this schema have different origins, so test any acceptance threshold with the provider, API, and task you intend to use.

Copy link to headingFrequently asked questions

Copy link to headingDoes this benchmark GPT-6.1 Sol?

No. These examples use Sol to judge supplied application data. Benchmarking those judgments requires expected answers and a process for comparing the model's results with them.

Copy link to headingDoes Sol return a confidence score for every evaluation answer?

No. This Responses example produces the fields defined in the application schema, without native probability distributions. Its unsupported-claim probability is a prompted estimate that needs testing before use as a decision threshold.

Copy link to headingCan I disable reasoning for every GPT-6 model?

No. GPT-6 Luna supports none, but GPT-6.1 Sol and GPT-6 Astra require a reasoning effort of low or higher. These settings apply to the Responses examples; native OpenAI Decisions does not accept Responses reasoning options.

Copy link to headingCan I evaluate a reply with an older OpenAI model?

Yes, if the model supports structured output through the Responses API. Check its supported reasoning settings and compare its judgments against reviewed examples before using it in the workflow.

More Decision models articles

Ready to deploy?