Studio previewSelf-host Studio locally.Run it

Conversation evals in the consumer application

Updated 52 minutes ago • October 3, 2026

createEvals runs scripted conversations against an existing createAgent instance. A case sends messages, assesses the visible answer or interrupt, resumes expected interrupts, and grades each criterion using a separate judge. It executes real workflows, tools, and application logic.

The SDK exposes serializable suites, a manifest, progress events, and results. Mount createEvalRouteHandler in the consumer app to let Studio execute suites and persist their observations and grades. See Studio execution for the endpoint, database migration, and local Docker configuration.

What an eval tests

A case is one use-case scenario: messages, expected answers or human pauses, and the criteria that define success. A suite groups related cases. A run executes selected cases, optionally several times. A judge grades the semantic criteria using the captured conversation and execution evidence. A dataset in other tools often serves the same role as a collection of cases; here a case can contain an entire conversation rather than a single input/output pair.

These are scripted workflow evals. They do not automatically grade ordinary user sessions. Production telemetry helps you find failures and author regression cases; running those cases creates new executions. It does not replay a historical session with its original permissions, data and side effects.

Keep deterministic SDK contracts in unit and integration tests. Use eval criteria for behavior such as choosing the right resource, asking for clarification, and using retrieved facts faithfully. A passing synthetic fixture verifies the eval plumbing; verify representative scenarios through the real application's execution and permission path before relying on their behavioral results.

Smallest configuration

import { createEvals, createEvalJudge, defineSuite } from "kortyx"; import { agent } from "./agent"; import { judgeModel } from "./models"; // The app's configured ProviderModelRef. const suite = defineSuite({ id: "product-information", cases: [{ id: "all-three-prices", steps: [ { message: "List the three products in my catalog", expect: { type: "answer" } }, { message: "Give me the prices for all three of those products", expect: { type: "answer", criteria: [ "Gives all three prices and currencies accurately from successful tool results, attributed to the requested products; fails if the required product data is missing.", "Does not ask the user to choose just one product.", ], }, }, ], }], }); const evals = createEvals({ agent, suites: [suite], judge: createEvalJudge({ model: judgeModel }), }); const result = await evals.run({ suiteId: "product-information", repetitions: 3, concurrency: 1, onProgress: async (event) => { // Save or stream this public event in your application. console.log(event); }, }); console.log(result.status, result.counts);

This example grades answers against the workflow's emitted tool results, so it needs no separate reference loader. Enable tool emission as shown below. Stable fixtures make behavior changes meaningful; changing live data can change the expected answer. An unexpected picker fails before calling the judge. Every configured criterion must pass. Without criteria, only the interaction type, optional interrupt schema and any required structured outputs are checked. Reference facts are optional; add them only when the scenario needs an independent expected result beyond the tool evidence.

Required structured outputs

Declare visible output contracts on the step that must produce them:

const suite = defineSuite({ id: "catalog-output", cases: [{ id: "list-and-summary", steps: [{ message: "Show the blue backpacks with prices and a short summary", expect: { type: "answer", outputs: [ { schemaId: "app.product-list" }, { schemaId: "app.product-summary", schemaVersion: "1" }, ], criteria: ["Prices match successful tool results for the requested products."], }, }], }], });

Every listed contract must have a completed visible output in that step. schemaVersion is optional: omitting it accepts any version; setting it requires an exact match. Multiple outputs can appear in one assistant response. Extra outputs are allowed, and this array does not assert their count or order. The registered workflow's output contract validates the payload; eval criteria judge its meaning and faithfulness to emitted tool results.

The runner checks the current finalized observation.structured envelopes before calling either an app or Studio judge. Partial outputs, invalidated outputs, outputs from earlier steps, tool results and IDs embedded inside a payload do not satisfy a requirement. A missing contract fails the case with the missing ID and version in the reason, skips semantic judging for that step and stops subsequent steps. Cleanup still runs. Output-only expectations need no LLM judge in local SDK runs; Studio runs still use its explicit judge selection flow.

Requirements appear in Studio's conversation plan and case evaluation, and are saved with the run's suite snapshot. The same JSON shape works for a suite authored in Studio, SDK execution and CLI-triggered runs.

outputs can also accompany an interrupt expectation when the workflow must show data before requesting a response. The top-level schemaId and schemaVersion on an interrupt expectation describe its single pending request; outputs describes completed visible outputs. Use another step to resume it. The default executor supports sequential interrupts, not simultaneous pending requests.

A custom executor must put the current step's finalized public envelopes in observation.structured, for example { schemaId: "app.product-list", schemaVersion: "2", status: "done", data: { products: [] } }. Return partials with status: "streaming", and remove invalidated entries. Do not copy earlier turns or tool results into this array. The native executor performs this accumulation automatically from agent.streamChat.

Execution evidence from existing tool outputs

For evals that check tool activity or ground answers in retrieved data, enable toolExecution.emit: true on every relevant workflow useReason call:

const result = await useReason({ model, input, tools, toolExecution: { emit: true }, });

The runner captures the existing public stream directly. Each step records ordered observation.events: tool names, call IDs, arguments, results and errors, alongside text, structured output, interrupt requests and workflow transitions. Adjacent text deltas from the same source are joined. The judge receives the current observation and all previous steps, including tool results retrieved before an interrupt. This works from the SDK, Studio and CLI without querying Studio telemetry.

A criterion can say: “Answers for each requested product with the price and currency returned by its successful tool call; explicitly states when price data is missing.” No separate reference loader is needed for this check. The judge checks faithfulness to the data the workflow actually received and can explain which call or step caused a failure. It accepts any valid execution path that satisfies the criterion. It cannot establish whether the external service itself is correct; use fixed reference facts when you also need to check against an independent expected result.

Tool emission defaults to off. Evals of answers or interrupts can still run without tool evidence, but criteria that require retrieved facts need emitted results or references. The eval runner cannot switch emission on for application workflows. Silent calls and private runtime state are outside the captured boundary. A custom executor must populate observation.events itself when it wants execution evidence. Only emit tool arguments and results suitable for the judge and stored eval records; they may contain application data, and arbitrary tool payloads are not automatically redacted.

Choose where judging runs

Workflow execution stays in the consumer application. Studio selects the judge for each run, defaulting to Studio judge, even when the app provides a code judge. The app needs no Studio model key or grading connection.

For Studio execution, no code judge is required:

import { createEvals } from "kortyx"; const evals = createEvals({ agent, suites: [suite], execute });

Provide a judge to enable direct local SDK runs and Studio's App judge option:

import { createEvals, createEvalJudge } from "kortyx"; const evals = createEvals({ agent, suites: [suite], execute, judge: createEvalJudge({ model: judgeModel }), }); // Uses the code judge. Studio is not required. const result = await evals.run({ suiteId: suite.id });

The app judge accepts the same provider references as your workflows. For example, with @kortyx/openrouter installed and OPENROUTER_API_KEY configured on the application server:

import { openrouter } from "@kortyx/openrouter"; import { createEvalJudge } from "kortyx"; const judge = createEvalJudge({ model: openrouter("openai/gpt-4o"), id: "product-information-judge", version: "1", });

Choose a provider route that supports structured JSON verdicts. See the OpenRouter provider for explicit provider settings. This is the App judge path; Studio judge credentials and its OpenRouter-compatible endpoint are configured separately on the Studio API.

A direct local run with semantic criteria and no code judge fails before setup or execution. Construction and suite discovery still work without a judge. Runs with no semantic criteria can execute locally using only interaction checks.

Local terminal runner

Export your createEvals instance as evals (or default) from a dedicated module. Run it directly without Studio:

pnpm exec kortyx evals run --entry ./src/evals/index.ts pnpm exec kortyx evals run --entry ./src/evals/index.ts --suite product-ambiguity --json

Add "eval": "kortyx evals run --entry ./src/evals/index.ts" to package.json for pnpm eval. The terminal report shows live progress, case verdicts, timings, and failed criteria with reasons/evidence. --json emits full SDK results for automation. All passes exit 0, failures/errors exit 1, cancellation exits 130. Your existing setup, executor, responders, cleanup and code judge are reused. No Studio model or execution credentials are needed. See CLI commands for filters, environment loading and report options.

When Studio judge is selected, the SDK bypasses the code judge, executes the scripted conversation, resolves declared interrupts and runs cleanup. Interaction mismatches and execution errors still stop the case. Successfully executed steps with semantic criteria have status: "ungraded" and no criterion verdicts. Studio then grades the captured evidence, records reasons and verdicts, and finishes the run. A case is not counted as passed while awaiting evaluation. Semantic grading happens after execution, so a semantic failure does not stop subsequent scripted steps in this mode. App judge retains per-step grading and stop-on-failure.

Studio's run drawer offers both available locations. The CLI defaults to Studio judging; pass --judge app to select the code judge. A missing judge or incompatible consumer SDK prevents starting that selection; there is no automatic fallback. The Studio backend pins the selected judge's ID and version when enqueuing. A changed judge causes an error before execution. Saved results include the identity and judging location for comparison.

Only captured public evidence is graded: criterion, input, observation, resolved reference when configured, and previous steps. Auth tokens and private runtime continuations stay in the app. Emitted tool payloads must be suitable for the judge and stored records. See the Studio execution guide for backend model configuration. Editing evaluator definitions in the browser and regrading saved runs are outside this release.

For advanced direct SDK use, createStudioEvalJudge({ url, apiKey, environment }) is also available as an explicit code judge. It calls the Studio API and requires a project key with studio:read and eval:run; this is not needed for ordinary Studio-triggered runs.

App-owned permissions and human responses

Keep suites as plain data and the application's authorization code in createEvals. No parameter schema or reference loader is needed for this example. The prompts name the scenario directly; the judge uses emitted tool results as evidence.

import { createEvals, defineSuite } from "kortyx"; export const productSuite = defineSuite({ id: "product-ambiguity", name: "Product information", cases: [{ id: "choose-blue", name: "Clarify an ambiguous product, then describe it", steps: [ { message: "Tell me about the Travel Backpack", expect: { type: "interrupt", schemaId: "app.product-picker", schemaVersion: "1", criteria: [ "Offers both the blue and black backpacks returned by successful tool calls as distinct choices, without selecting one on the user's behalf.", ], }, }, { resume: { using: "chooseBlue" }, expect: { type: "answer", criteria: [ "Describes the selected blue backpack faithfully using its successful tool result.", ], }, }, ], }], }); const evals = createEvals({ agent, suites: [productSuite], setup: async ({ signal }) => { // App-owned helper: obtains the configured test actor's credentials. const accessToken = await obtainTestUserToken(signal); return { accessToken }; }, execute: async ({ prepared, run }) => { // App-owned helper: the same identity and permission scope as normal requests. // Await run() inside the scope so request-bound tools retain authorization. return withAuthorizedAgentRequest(prepared.accessToken, () => run()); }, responders: { chooseBlue: { schemaId: "app.product-picker", schemaVersion: "1", // App-owned helper: reads the actual emitted picker request and returns // { type: "value", value: <your contract's validated selection> }. respond: ({ interrupt }) => selectProductByColor(interrupt, "blue"), }, }, });

Configure the test actor and its accessible data in the app. This scenario assumes both roles exist and are visible to that actor. If permissions or data expose only one role, author a separate case with an answer expectation. The responder fails if it cannot find the intended role in the actual choices; it must not invent a resource ID. The application helpers above are not imports from Kortyx.

setup runs once per case repetition and returns private prepared state. The runner passes it to execute and named handlers; it is not sent to Studio. Setup can acquire an Auth0 token, prepare fixtures or bind a test identity. Kortyx does not implement an Auth0 login flow. execute adapts each scripted command to the app's normal execution boundary, including its permission and request context. Reuse the same application helper that your chat/API route uses so the eval does not bypass caller restrictions.

An identity-provider token alone does not configure permissions inside Kortyx. The app must pass it through its existing authorization path. With no execute, the runner calls agent.streamChat directly without injecting setup state or headers. run({ context }) consumes that same native stream; it adds the current command and handles native resume metadata. It does not call an HTTP endpoint.

For an app-owned HTTP or typed execution API, replace execute with an adapter that handles command (message, resume, or cancel), history, sessionId, signal, and its private continuation. Return { observation, continuation? }. An interrupt observation needs a continuation so the next command can resume or cancel it. Consume execution to completion or the intended human pause before returning. HTTP response parsing, API auth, and mapping app-owned UI actions remain the adapter's responsibility.

Native resume responses use { type: "text", text }, { type: "select", ids }, { type: "value", value }, or { type: "cancel" }. The value envelope contains the custom contract's actual response. Returning the wrong contract value fails during native resume validation. A named responder with a schema descriptor is checked before it consumes the request. Named responders and references can also receive JSON arguments through { using: "handlerName", params: { ... } }.

teardown releases app resources; it need not sign out. It runs after successful setup, including after failed expectations, cancellation, and errors. Setup must clean up partially created resources if it throws before returning.

Observable boundary and results

The default executor assesses finalized visible text, streamed execution events, structured data, and one pending interrupt per step. Structured entries retain their public stream envelope, including dataType, data, and status. The executor excludes done.data, transition payloads and arbitrary interrupt transport metadata. Raw workflow state, silent tools, private setup state, and resume tokens are not copied into eval results or judge input. Run/checkpoint IDs connect observations to diagnostics when the agent supplies them. Custom executor observations are validated.

References are independent facts from fixtures or application reads. An inline JSON reference works too. Objects containing only using and optional params are reserved for named reference handlers. The runner snapshots the resolved reference alongside each observation. Do not put secrets in case params, messages, emitted tool arguments/results, reference facts, custom observations, or responder values: these fields are intentionally visible and serializable.

createEvalJudge invokes the separately configured provider model once per criterion, requests a structured JSON verdict, and validates it. A verdict has passed, reason, and evidence. A malformed, refused, or truncated verdict is an error, not a failed behavior score. The model's judgments still need human calibration against representative passes and failures. You can supply an EvalJudge with your own grade implementation instead.

Each run records the suite snapshot and revision hash, judge ID/version, unique session IDs per repetition, observations, references, criterion verdicts, counts, timing, and errors by phase. passed, failed, error, and cancelled remain distinct. Interaction mismatches and execution errors stop the case. App judging also stops after a semantic failure; Studio judging grades after the conversation has executed. This is not a complete reproducibility record of model versions, prompts, app code, and external data; retain those deployment details in your app's run record.

Raw exception messages are excluded from public results because they can contain credentials. Log detailed failures privately in your app's hooks or existing runtime diagnostics. Progress events are snapshots; listener failures mark the run as a reporting error without skipping case cleanup.

Lifecycle limits in this slice

  • Sequential interrupts, including child workflows and custom contracts, use native resume execution. Remaining observed interrupts are cancelled during cleanup, including when the case intentionally ends on an interrupt.
  • Simultaneous pending interrupts are reported as an execution error by the default executor, which attempts to cancel every observed request. Provide a custom executor for a domain-specific parallel interaction protocol.
  • Early completeResponse/background continuation requires a custom executor. The default runner awaits attached execution completion and rejects a response stream that does not expose a completed workflow or human pause.
  • Case and cleanup deadlines use cooperative abort signals. Arbitrary app callbacks must honor their signals; the runner cannot forcibly terminate JavaScript. It waits for work to settle before teardown to avoid racing cleanup against an execution that is still using its fixtures.
  • Sessions are isolated, but databases and external side effects are only as isolated as the app's setup/teardown. Native cancellation does not delete persisted session records. Internal effects need explicit app checks; SDK deterministic behavior belongs in normal integration tests.

evals.listSuites() returns editable copies. evals.describe() returns schema version 1, suites, named handlers, optional parameter JSON Schema, and judge identity. Neither invokes setup. createEvalRouteHandler exposes this manifest and runs to the Studio control plane. See Studio execution for authenticated transport, persistent runs, and local Docker setup. Saved-run comparisons are available in Studio. Prompt version pinning remains subsequent work.

Optional typed case parameters

Most suites need no case parameters or schema: write prompts and expectations directly. When app setup needs configurable actor or fixture identifiers, the SDK also supports public JSON case parameters.

defineSuite supplies the suite structure, so you do not need a separate EvalSuite annotation. Use defineSuite<TParams> to check case parameters while authoring. Keep the Zod schema in createEvals for runtime validation.

import { createEvals, defineSuite } from "kortyx"; import { z } from "zod"; const paramsSchema = z.object({ productIds: z.array(z.string()).min(1).max(3), }).strict(); type ProductParams = z.input<typeof paramsSchema>; export const authorizedProductsSuite = defineSuite<ProductParams>({ id: "authorized-products", cases: [{ id: "product-description", params: { productIds: ["blue-backpack"] }, steps: [{ message: "Describe the blue Travel Backpack", expect: { type: "answer" }, }], }], }); const evals = createEvals({ agent, paramsSchema, suites: [authorizedProductsSuite], setup: async ({ params }) => { // params.productIds is string[] here and in the suite definition above. return prepareProductFixture(params.productIds); }, });

Missing required parameters, misspelled fields, and incompatible values such as productIds: [123] are TypeScript errors. Constraints such as minimum array length are checked at runtime. Derive TParams using z.input<typeof paramsSchema> so defaults and transforms are applied once by the runner before setup receives the schema's output type. The suite remains plain data. You can also use satisfies EvalSuite<ProductParams> without the helper. Without a type argument, defineSuite checks the standard suite structure.

The parameters above belong to suite.cases[].params. They are neither workflow output nor automatically discovered tool context. One paramsSchema validates all cases registered on that createEvals instance. For several actor or fixture shapes, use an app-owned discriminated union; the same endpoint can execute them without creating a separate endpoint per user. Keep credentials in setup rather than public parameters.

defineSuite also works in JavaScript without a generic type argument. JSON suites can be validated with parseEvalSuite(json) before registration. Named responders and references must already exist in the app; serialized definitions cannot install executable callbacks. This release displays suite definitions in Studio; authoring or saving new suites in Studio is not implemented.