Studio previewSelf-host Studio locally.Run it

Run consumer evals from Studio

Updated 52 minutes ago • October 3, 2026

Studio can discover registered suites, select cases and repetitions, enqueue a run, observe each completed interaction, cancel execution, and reload saved results. The consumer app executes its own agent. Studio does not obtain end-user tokens, implement application permissions, or invoke the application's chat endpoint.

import { createEvals, createEvalRouteHandler } from "kortyx"; const evals = createEvals({ agent, suites, setup, // optional: authenticate the test actor or prepare fixtures execute, // bind the app's normal permission/tool context, then await run() responders, // optional: app handlers referenced by JSON resume steps teardown, // optional: release resources prepared for each case // No code judge required: Studio can grade the captured execution. }); const handleEvals = createEvalRouteHandler({ evals, serviceKey: process.env.EVAL_SERVICE_KEY!, }); // Mount handleEvals(Web Request) -> Web Response using your server framework.

GET returns the public manifest. POST accepts a suite ID, its SHA-256 revision, optional case IDs, repetitions, and concurrency. It streams validated NDJSON progress records and one final result. A changed suite is rejected before execution. The handler requires a bearer service key of at least 32 characters, bounds attempts and concurrent runs, and signals cooperative cancellation when its client disconnects.

For every variable's default, owning service, credential purpose and restart requirements, see the Configuration Reference.

Register applications on the Studio API server

Set KORTYX_EVAL_TARGETS_FILE to a secret-manager supplied JSON file:

[ { "id": "catalog", "name": "Catalog agent", "organizationId": "YOUR-ORGANIZATION-UUID", "projectId": "YOUR-PROJECT-UUID", "environment": "development", "url": "https://your-app.example/v1/evals", "serviceKey": "YOUR-APPLICATION-EVAL-SERVICE-KEY" } ]

The server fixes target URLs and credentials; the browser cannot submit an arbitrary URL or token. HTTPS is required unless allowInsecureHttp: true is explicitly set for a local target. Targets, history, details and cancellation are scoped to the Studio API key's organization and project. Execution also requires eval:run and an allowed project environment. Reading uses the existing studio:read scope. The current browser proxy supports local none and deployed basic Studio auth modes. Suite execution uses the existing server-side Studio key; it does not require a new login system. For the local Docker bootstrap, opt in with KORTYX_STUDIO_ENABLE_EVALS=1 and rerun bootstrap with the same stored Studio key. Execution is disabled by default.

Where application metadata comes from

The target file supplies the application ID, display name, endpoint and environment. They are operator-defined labels, not inferred from model calls or the agent's telemetry service name. The application's manifest supplies the suite IDs, case definitions, revisions, named handlers and optional code judge.

Studio copies targetId, targetName and environment into each saved run. The application filter combines currently configured targets with labels from saved history. Removing a target stops discovery and new execution, but its historical runs and application filter entry remain. Suite discovery shows currently registered suites; saved runs retain the definition used when they started.

Use stable target IDs. Renaming a target changes its label for new runs; existing records keep the saved label. The filter groups by target ID and displays a label from the current configuration or saved history. Two environments can be registered as separate targets with distinct IDs and clear names.

Credentials have separate jobs

CredentialWhere it is configuredWhat it authorizes
Project Studio keyStudio server or CLI connectionProject discovery/history with studio:read; execution/cancellation with eval:run
Application eval service keyConsumer route and Studio API target fileStudio API access to that app's eval endpoint
Test actor credentialsConsumer application's server configuration/setupThe app's normal user identity, roles and data access
Judge provider keyStudio API backend for Studio judging; consumer app for App judgingModel requests for semantic grading

The service key authorizes starting evals; it does not impersonate an application user. The consumer's setup and execute callbacks bind that user through the same permission path used by ordinary application requests.

Studio judge configuration

Studio-triggered runs default to Studio judge. Configure its model and provider key on the Studio API backend. A code judge provided to createEvals makes App judge available as an explicit selection; Studio still defaults to its own judge. If hosted judging is not configured, the run drawer explains what is missing and allows selecting an available App judge. It never silently changes selection.

The app executes the complete scripted scenario and resolves interrupts. With Studio judging, it returns captured evidence awaiting evaluation, then the Studio worker grades that evidence and stores the final results. App judging runs inside the app at each step. The browser never calls a model directly. The app needs no Studio grading key or model configuration for the hosted path.

Enable the backend judge on the API service:

KORTYX_EVAL_JUDGE_MODEL=gpt-5.4-mini KORTYX_EVAL_JUDGE_API_KEY=<server-owned provider key> # Optional overrides: # KORTYX_EVAL_JUDGE_ID=studio/my-judge # KORTYX_EVAL_JUDGE_VERSION=kortyx-rubric-v3 # KORTYX_EVAL_JUDGE_API=responses # KORTYX_EVAL_JUDGE_BASE_URL=https://api.openai.com/v1

The default server adapter uses Kortyx's OpenAI provider and its Responses API. Set KORTYX_EVAL_JUDGE_API=chat-completions for compatible Chat Completions endpoints. A custom HTTPS endpoint must support the selected API and structured JSON verdicts. Custom API hosts can instead inject any EvalJudge through createApiApp({ ..., evalJudge }), including a judge built with another Kortyx provider. Leaving the model unset disables hosted judging; app-side judging and ordinary Studio execution remain available.

For OpenRouter, use its model slug and server-side API key:

KORTYX_EVAL_JUDGE_BASE_URL=https://openrouter.ai/api/v1 KORTYX_EVAL_JUDGE_API=chat-completions KORTYX_EVAL_JUDGE_MODEL=openai/gpt-4o KORTYX_EVAL_JUDGE_API_KEY=<OpenRouter API key> KORTYX_EVAL_JUDGE_ID=studio/openrouter/openai/gpt-4o

Choose a model/provider combination that supports structured outputs. The OpenRouter chat endpoint uses Kortyx's native OpenRouter provider, including its reported billing usage. A live OpenRouter account is needed to verify the configured model and provider route.

Docker Compose and CLI-generated Studio stacks pass these variables only to the API container. For a CLI-managed stack, add them to the existing owner-only <studio-home>/.env and restart with kortyx studio start --home <studio-home>. CLI startup preserves these private settings. For other deployments, use the API service's normal secret configuration. Provider credentials are never returned by judge discovery or included in run results.

The browser and CLI use the existing project key with studio:read and eval:run to enqueue runs. The app's eval service key stays in server target configuration. Project and environment access is enforced on the Studio API. Provider credentials are never sent to the app or browser.

Suite discovery advertises a code judge when one is configured and the consumer's Studio judging capability. The run drawer stores its selection in the URL. The API pins the selected judge identity in the saved request; the worker rejects a changed Studio judge before starting the app. The consumer similarly rejects a changed code judge. Missing configuration, provider failures and malformed verdicts produce errors instead of passing scores. Captured cases display Awaiting evaluation until Studio has graded them. Cancellation covers execution and grading.

The optional /v1/studio/evals/judge endpoint remains available for explicit createStudioEvalJudge calls from code. It authenticates each request using a project Studio key, checks its environment, pins identity and bounds input size, concurrency and grading deadlines. Ordinary Studio runs grade within the worker and do not call this endpoint. Saved-run regrading and editable evaluator libraries are not included in this release.

Eval run costs

Run history and each run's case table show model costs. The run summary and case inspector provide a Workflow / Judge / Total breakdown. Workflow cost uses recorded generation telemetry for the attempt's session in the same project and environment, including child workflows, retries and resumed turns. Enable normal agent telemetry to record these calls. toolExecution.emit: true supplies tool evidence to the judge; it does not itself enable billing telemetry.

createEvalJudge captures provider usage separately from the generated verdict. Both app-owned and Studio-owned judges report it, including a paid call whose verdict fails validation. OpenRouter reports actual charges; BYOK uses upstream inference cost when supplied. Other model usage uses Studio's effective model rate cards. These calculated charges are marked estimated in the tooltip. No extra cost-related environment variable is required. Custom judges may report usage through EvalGradeInput.onUsage; without it their cost remains unknown.

A + after a displayed amount means a known subtotal, with more cost possible. A dash means unavailable, including absent telemetry or prices. Running or unfinished attempts stay partial; currencies are combined only when compatible. Historical workflow costs can be recovered from existing telemetry, but historical judge charges without recorded usage cannot be recovered. Costs cover recorded model calls, excluding infrastructure and external tool fees. They are not an independent reconciliation of a provider invoice. Cancellation can leave billing evidence incomplete when a remote call finishes after the executor disconnects.

Local kortyx evals run --entry … still executes without Studio and does not save a Studio record. Use the registered Studio execution endpoint to save runs and view their combined costs in Studio.

Persistence and execution lifecycle

Apply packages/telemetry-db/drizzle/0005_eval_runs.sql through the regular migration runner before starting the updated API. Each queued record stores its suite and revision, selection, target and environment. PostgreSQL row locking allows a worker replica to claim a record once. A lease and heartbeat track active execution. Progress is saved incrementally; the final result contains public observations (including emitted execution events), expected behavior, optional references, grader version, reasons, and evidence. Studio receives scoped SSE change notifications backed by PostgreSQL LISTEN/NOTIFY, then reads the saved result. Runs, case drawers and history update on changes; completed runs also receive late workflow billing updates. Live mode is on by default for eval views and can be paused with ?live=false. The connection pauses while the tab is hidden or offline, reconciles on reconnect, and uses a 30–60 second refresh fallback during an outage. There is no periodic polling while connected. A browser reload does not terminate the run. Caller-visible observations may contain business data and require the same access and retention policy as ordinary Studio session data.

A consumer disconnect, timeout, or worker crash is an execution error rather than a failed behavioral grade. An expired lease is marked unknown/error and is never automatically retried: replaying application workflows can repeat side effects. Operators should inspect consumer execution before explicitly starting another run. Cancellation wins the final-storage race when its request is already saved. The worker actively closes an aborted response reader to signal the consumer; custom setup, tools and cleanup must honor the signal. A worker shutdown does not promise rollback or prove remote cleanup.

Studio supports saved-run comparison, including per-case outcomes, criterion verdicts, repeated attempts and workflow inspection through stacked drawers. Comparison describes observed differences. Its context notice identifies missing actor/data/prompt identity; it does not establish why an outcome changed.

This release supports suites authored in app code or JSON, sequential interactions, and LLM grading of public answers, interrupt requests and emitted tool evidence. Tool-based grading requires toolExecution.emit: true in the workflow. The judge uses the existing stream, including results from earlier steps before a resume. See the conversation eval guide for the capture boundary. Studio suite editing, prompt-variant injection, scheduled runs, retention controls, and an executor protocol for simultaneous interrupts or durable background work remain subsequent work. Prompt changes can be checked by rerunning the same suite against the changed consumer; prompt/version pinning is not implemented.

Use the CLI commands and post-deployment CI example to discover, start, inspect and cancel runs from the same Studio API.

Use the interface

Open Evals → Suites to browse the applications and cases registered by your consumer. Select a suite to inspect its definition in a drawer, or open it as a full page. Choose cases, repetitions and concurrency, then start the run.

Evals → Runs lists saved runs. A run shows live progress and final case outcomes. Open a case to read evaluation reasons first, then expand conversation and debugging details. Inspect workflow opens the existing workflow drawer on top; closing it returns to the case. Page URLs, filters and selections survive reloads and browser Back/Forward.

Use Compare on a saved run and select a baseline from the same target/suite. The candidate is the run you are inspecting. The baseline is another saved run you choose as a reference; it is not an expected answer or a newly executed workflow. Comparison reads their saved grades without calling a judge or rerunning either application.

Cases are matched by case ID. For comparable cases, Studio compares the share of passed attempts, including repetitions:

LabelMeaning
ImprovedCandidate has a higher pass rate
RegressedCandidate has a lower pass rate
UnchangedBoth have the same pass rate; their outputs can still differ
IncompleteAn attempt is missing, ungraded, errored, cancelled or otherwise lacks a completed result
Context changed / not comparableTarget, environment, case definition, judge identity, recorded input or reference differs

For example, 1 of 3 passed in the baseline and 3 of 3 in the candidate is Improved. A provider error is Incomplete, not a behavioral regression. Open a comparison case to inspect both sides' attempts, outputs, criteria and workflow links.

Actor identity, external data snapshots and prompt versions are not automatically pinned. Tool results can change even when the recorded inputs are identical. Treat comparison as observed pass-rate differences, not proof that a prompt change caused them. Rerun the same suite with stable actor/data and review the actual evidence when testing a prompt change.

Container configuration

Mount the private target file on every Studio API replica and set the variable on the API service, using the container path. The browser Studio service does not need the file or application service key. For Docker Compose:

services: api: environment: KORTYX_EVAL_TARGETS_FILE: /run/secrets/kortyx-eval-targets.json volumes: - ./eval-targets.json:/run/secrets/kortyx-eval-targets.json:ro db-init: environment: KORTYX_STUDIO_ENABLE_EVALS: "1"

Use this as a deployment override alongside the normal Compose file. For a CLI-managed local installation, keep it outside generated state and apply it when starting/recreating services: studio start regenerates its base Compose file. Retain the existing environment file and project name, and rerun db-init to update the existing Studio key scopes before recreating the API.

The bootstrap output provides the organization and project IDs for target registration. Add an allowed project environment before using it in a target. A container cannot reach a host application at localhost; Docker Desktop typically uses host.docker.internal, or use an application on the same Docker network. Local HTTP targets need allowInsecureHttp: true; remote targets use HTTPS.

Targets are loaded when the API starts. Restart the API after changing the target file. Missing targets leave Suites empty; an unreachable endpoint shows the target unavailable. Run controls require eval:run on the existing key. Read-only keys can still inspect saved results. An endpoint's service key is a separate credential from the Studio key; neither contains the test user's token.

See Conversation Evals for suite design, typed setup/responders, Auth0 integration and grading boundaries.

Required structured responses

Conversation steps can declare expect.outputs as an array of required output contracts. Studio displays each ID and either the exact required version or “Any version” in the suite definition and case evaluation. The consumer checks that every contract has a completed visible output before either judge runs. Partial streams, invalidated outputs and outputs from previous steps do not satisfy the requirement. Missing contracts fail the step with an explicit reason. Open the conversation debugging section to inspect the recorded envelopes. Use pass criteria to assess payload meaning; the output requirements check contract presence. See the conversation guide for suite authoring and custom executor support.