# Run consumer evals from Studio

Studio can discover registered suites, select cases and repetitions, enqueue a run,
observe each completed interaction, cancel execution, and reload saved results.
The consumer app executes its own agent. Studio does not obtain end-user tokens,
implement application permissions, or invoke the application's chat endpoint.

```ts
import { createEvals, createEvalRouteHandler } from "kortyx";

const evals = createEvals({
  agent,
  suites,
  setup,       // optional: authenticate the test actor or prepare fixtures
  execute,     // bind the app's normal permission/tool context, then await run()
  responders,  // optional: app handlers referenced by JSON resume steps
  teardown,    // optional: release resources prepared for each case
  // No code judge required: Studio can grade the captured execution.
});
const handleEvals = createEvalRouteHandler({
  evals,
  serviceKey: process.env.EVAL_SERVICE_KEY!,
});
// Mount handleEvals(Web Request) -> Web Response using your server framework.
```

`GET` returns the public manifest. `POST` accepts a suite ID, its SHA-256 revision,
optional case IDs, repetitions, and concurrency. It streams validated NDJSON
progress records and one final result. A changed suite is rejected before execution.
The handler requires a bearer service key of at least 32 characters, bounds
attempts and concurrent runs, and signals cooperative cancellation when its
client disconnects.

For every variable's default, owning service, credential purpose and restart
requirements, see the [Configuration Reference](./08-configuration-reference.md).

## Register applications on the Studio API server

Set `KORTYX_EVAL_TARGETS_FILE` to a secret-manager supplied JSON file:

```json
[
  {
    "id": "catalog",
    "name": "Catalog agent",
    "organizationId": "YOUR-ORGANIZATION-UUID",
    "projectId": "YOUR-PROJECT-UUID",
    "environment": "development",
    "url": "https://your-app.example/v1/evals",
    "serviceKey": "YOUR-APPLICATION-EVAL-SERVICE-KEY"
  }
]
```

The server fixes target URLs and credentials; the browser cannot submit an arbitrary
URL or token. HTTPS is required unless `allowInsecureHttp: true` is explicitly set
for a local target. Targets, history, details and cancellation are scoped to the
Studio API key's organization and project. Execution also requires `eval:run` and
an allowed project environment. Reading uses the existing `studio:read` scope. The current browser proxy supports
local `none` and deployed `basic` Studio auth modes. Suite execution uses the
existing server-side Studio key; it does not require a new login system.
For the local Docker bootstrap, opt in with `KORTYX_STUDIO_ENABLE_EVALS=1` and rerun
bootstrap with the same stored Studio key. Execution is disabled by default.

### Where application metadata comes from

The target file supplies the application **ID**, display **name**, endpoint and
**environment**. They are operator-defined labels, not inferred from model calls
or the agent's telemetry service name. The application's manifest supplies the
suite IDs, case definitions, revisions, named handlers and optional code judge.

Studio copies `targetId`, `targetName` and environment into each saved run. The
application filter combines currently configured targets with labels from saved
history. Removing a target stops discovery and new execution, but its historical
runs and application filter entry remain. Suite discovery shows currently
registered suites; saved runs retain the definition used when they started.

Use stable target IDs. Renaming a target changes its label for new runs; existing
records keep the saved label. The filter groups by target ID and displays a label
from the current configuration or saved history. Two environments can be
registered as separate targets with distinct IDs and clear names.

### Credentials have separate jobs

| Credential | Where it is configured | What it authorizes |
| --- | --- | --- |
| Project Studio key | Studio server or CLI connection | Project discovery/history with `studio:read`; execution/cancellation with `eval:run` |
| Application eval service key | Consumer route and Studio API target file | Studio API access to that app's eval endpoint |
| Test actor credentials | Consumer application's server configuration/setup | The app's normal user identity, roles and data access |
| Judge provider key | Studio API backend for Studio judging; consumer app for App judging | Model requests for semantic grading |

The service key authorizes starting evals; it does not impersonate an application
user. The consumer's setup and execute callbacks bind that user through the same
permission path used by ordinary application requests.

## Studio judge configuration

Studio-triggered runs default to **Studio judge**. Configure its model and provider
key on the Studio **API backend**. A code judge provided to `createEvals` makes
**App judge** available as an explicit selection; Studio still defaults to its own
judge. If hosted judging is not configured, the run drawer explains what is missing
and allows selecting an available App judge. It never silently changes selection.

The app executes the complete scripted scenario and resolves interrupts. With
Studio judging, it returns captured evidence awaiting evaluation, then the Studio
worker grades that evidence and stores the final results. App judging runs inside
the app at each step. The browser never calls a model directly. The app needs no
Studio grading key or model configuration for the hosted path.

Enable the backend judge on the **API service**:

```dotenv
KORTYX_EVAL_JUDGE_MODEL=gpt-5.4-mini
KORTYX_EVAL_JUDGE_API_KEY=<server-owned provider key>
# Optional overrides:
# KORTYX_EVAL_JUDGE_ID=studio/my-judge
# KORTYX_EVAL_JUDGE_VERSION=kortyx-rubric-v3
# KORTYX_EVAL_JUDGE_API=responses
# KORTYX_EVAL_JUDGE_BASE_URL=https://api.openai.com/v1
```

The default server adapter uses Kortyx's OpenAI provider and its Responses API.
Set `KORTYX_EVAL_JUDGE_API=chat-completions` for compatible Chat Completions
endpoints. A custom HTTPS endpoint must support the selected API and structured
JSON verdicts.
Custom API hosts can instead inject any `EvalJudge` through
`createApiApp({ ..., evalJudge })`, including a judge built with another Kortyx
provider. Leaving the model unset disables hosted judging; app-side judging and
ordinary Studio execution remain available.

For OpenRouter, use its model slug and server-side API key:

```dotenv
KORTYX_EVAL_JUDGE_BASE_URL=https://openrouter.ai/api/v1
KORTYX_EVAL_JUDGE_API=chat-completions
KORTYX_EVAL_JUDGE_MODEL=openai/gpt-4o
KORTYX_EVAL_JUDGE_API_KEY=<OpenRouter API key>
KORTYX_EVAL_JUDGE_ID=studio/openrouter/openai/gpt-4o
```

Choose a model/provider combination that supports [structured outputs](https://openrouter.ai/docs/guides/features/structured-outputs).
The OpenRouter chat endpoint uses Kortyx's native OpenRouter provider, including
its reported billing usage. A live OpenRouter account is needed to verify the
configured model and provider route.

Docker Compose and CLI-generated Studio stacks pass these variables only to the
API container. For a CLI-managed stack, add them to the existing owner-only
`<studio-home>/.env` and restart with `kortyx studio start --home <studio-home>`.
CLI startup preserves these private settings. For other deployments, use the
API service's normal secret configuration. Provider credentials are never
returned by judge discovery or included in run results.

The browser and CLI use the existing project key with `studio:read` and `eval:run`
to enqueue runs. The app's eval service key stays in server target configuration.
Project and environment access is enforced on the Studio API. Provider credentials
are never sent to the app or browser.

Suite discovery advertises a code judge when one is configured and the consumer's
Studio judging capability. The run drawer stores its selection in the URL. The API
pins the selected judge identity in the saved request; the worker rejects a changed
Studio judge before starting the app. The consumer similarly rejects a changed
code judge. Missing configuration, provider failures and malformed verdicts produce
errors instead of passing scores. Captured cases display **Awaiting evaluation**
until Studio has graded them. Cancellation covers execution and grading.

The optional `/v1/studio/evals/judge` endpoint remains available for explicit
`createStudioEvalJudge` calls from code. It authenticates each request using a
project Studio key, checks its environment, pins identity and bounds input size,
concurrency and grading deadlines. Ordinary Studio runs grade within the worker
and do not call this endpoint. Saved-run regrading and editable evaluator libraries
are not included in this release.

## Eval run costs

Run history and each run's case table show model costs. The run summary and
case inspector provide a **Workflow / Judge / Total** breakdown. Workflow cost
uses recorded generation telemetry for the attempt's session in the same project
and environment, including child workflows, retries and resumed turns. Enable
normal agent telemetry to record these calls. `toolExecution.emit: true` supplies
tool evidence to the judge; it does not itself enable billing telemetry.

`createEvalJudge` captures provider usage separately from the generated verdict.
Both app-owned and Studio-owned judges report it, including a paid call whose
verdict fails validation. OpenRouter reports actual charges; BYOK uses upstream
inference cost when supplied. Other model usage uses Studio's effective model
rate cards. These calculated charges are marked **estimated** in the tooltip.
No extra cost-related environment variable is required. Custom judges may report
usage through `EvalGradeInput.onUsage`; without it their cost remains unknown.

A `+` after a displayed amount means a known subtotal, with more cost possible.
A dash means unavailable, including absent telemetry or prices. Running or
unfinished attempts stay partial; currencies are combined only when compatible.
Historical workflow costs can be recovered from existing telemetry, but historical
judge charges without recorded usage cannot be recovered. Costs cover recorded
model calls, excluding infrastructure and external tool fees. They are not an
independent reconciliation of a provider invoice. Cancellation can leave billing
evidence incomplete when a remote call finishes after the executor disconnects.

Local `kortyx evals run --entry …` still executes without Studio and does not save
a Studio record. Use the registered Studio execution endpoint to save runs and
view their combined costs in Studio.

## Persistence and execution lifecycle

Apply `packages/telemetry-db/drizzle/0005_eval_runs.sql` through the regular
migration runner before starting the updated API. Each queued record stores its
suite and revision, selection, target and environment. PostgreSQL row locking
allows a worker replica to claim a record once. A lease and heartbeat track active
execution. Progress is saved incrementally; the final result contains public
observations (including emitted execution events), expected behavior, optional
references, grader version, reasons,
and evidence. Studio receives scoped SSE change notifications backed by PostgreSQL
`LISTEN`/`NOTIFY`, then reads the saved result. Runs, case drawers and history
update on changes; completed runs also receive late workflow billing updates.
Live mode is on by default for eval views and can be paused with `?live=false`.
The connection pauses while the tab is hidden or offline, reconciles on reconnect,
and uses a 30–60 second refresh fallback during an outage. There is no periodic
polling while connected. A browser reload does not terminate the run.
Caller-visible observations may contain business data and require the same access
and retention policy as ordinary Studio session data.

A consumer disconnect, timeout, or worker crash is an execution error rather than
a failed behavioral grade. An expired lease is marked unknown/error and is never
automatically retried: replaying application workflows can repeat side effects.
Operators should inspect consumer execution before explicitly starting another run.
Cancellation wins the final-storage race when its request is already saved.
The worker actively closes an aborted response reader to signal the consumer; custom setup, tools and cleanup must honor the
signal. A worker shutdown does not promise rollback or prove remote cleanup.

Studio supports saved-run comparison, including per-case outcomes, criterion
verdicts, repeated attempts and workflow inspection through stacked drawers.
Comparison describes observed differences. Its context notice identifies missing
actor/data/prompt identity; it does not establish why an outcome changed.

This release supports suites authored in app code or JSON, sequential interactions,
and LLM grading of public answers, interrupt requests and emitted tool evidence.
Tool-based grading requires `toolExecution.emit: true` in the workflow. The judge
uses the existing stream, including results from earlier steps before a resume.
See the conversation eval guide for the capture boundary. Studio suite editing,
prompt-variant injection, scheduled runs, retention controls,
and an executor protocol for simultaneous interrupts or durable background work
remain subsequent work. Prompt changes can be checked by rerunning the same
suite against the changed consumer; prompt/version pinning is not implemented.

Use the [CLI commands and post-deployment CI example](./04-cli-commands.md#eval-suites-and-post-deployment-ci) to discover,
start, inspect and cancel runs from the same Studio API.

## Use the interface

Open **Evals → Suites** to browse the applications and cases registered by your
consumer. Select a suite to inspect its definition in a drawer, or open it as a
full page. Choose cases, repetitions and concurrency, then start the run.

**Evals → Runs** lists saved runs. A run shows live progress and final case
outcomes. Open a case to read evaluation reasons first, then expand conversation
and debugging details. **Inspect workflow** opens the existing workflow drawer
on top; closing it returns to the case. Page URLs, filters and selections survive
reloads and browser Back/Forward.

Use **Compare** on a saved run and select a baseline from the same target/suite.
The **candidate** is the run you are inspecting. The **baseline** is another
saved run you choose as a reference; it is not an expected answer or a newly
executed workflow. Comparison reads their saved grades without calling a judge
or rerunning either application.

Cases are matched by case ID. For comparable cases, Studio compares the share
of passed attempts, including repetitions:

| Label | Meaning |
| --- | --- |
| Improved | Candidate has a higher pass rate |
| Regressed | Candidate has a lower pass rate |
| Unchanged | Both have the same pass rate; their outputs can still differ |
| Incomplete | An attempt is missing, ungraded, errored, cancelled or otherwise lacks a completed result |
| Context changed / not comparable | Target, environment, case definition, judge identity, recorded input or reference differs |

For example, 1 of 3 passed in the baseline and 3 of 3 in the candidate is
**Improved**. A provider error is **Incomplete**, not a behavioral regression.
Open a comparison case to inspect both sides' attempts, outputs, criteria and
workflow links.

Actor identity, external data snapshots and prompt versions are not automatically
pinned. Tool results can change even when the recorded inputs are identical.
Treat comparison as observed pass-rate differences, not proof that a prompt
change caused them. Rerun the same suite with stable actor/data and review the
actual evidence when testing a prompt change.

## Container configuration

Mount the private target file on every Studio API replica and set the variable
on the API service, using the container path. The browser Studio service does
not need the file or application service key. For Docker Compose:

```yaml
services:
  api:
    environment:
      KORTYX_EVAL_TARGETS_FILE: /run/secrets/kortyx-eval-targets.json
    volumes:
      - ./eval-targets.json:/run/secrets/kortyx-eval-targets.json:ro
  db-init:
    environment:
      KORTYX_STUDIO_ENABLE_EVALS: "1"
```

Use this as a deployment override alongside the normal Compose file. For a
CLI-managed local installation, keep it outside generated state and apply it
when starting/recreating services: `studio start` regenerates its base Compose
file. Retain the existing environment file and project name, and rerun
`db-init` to update the existing Studio key scopes before recreating the API.

The bootstrap output provides the organization and project IDs for target
registration. Add an allowed project environment before using it in a target.
A container cannot reach a host application at `localhost`; Docker Desktop
typically uses `host.docker.internal`, or use an application on the same Docker
network. Local HTTP targets need `allowInsecureHttp: true`; remote targets use HTTPS.

Targets are loaded when the API starts. Restart the API after changing the
target file. Missing targets leave Suites empty; an unreachable endpoint shows
the target unavailable. Run controls require `eval:run` on the existing key.
Read-only keys can still inspect saved results. An endpoint's service key is a
separate credential from the Studio key; neither contains the test user's token.

See [Conversation Evals](../03-guides/10-conversation-evals.md) for suite design,
typed setup/responders, Auth0 integration and grading boundaries.

## Required structured responses

Conversation steps can declare `expect.outputs` as an array of required output
contracts. Studio displays each ID and either the exact required version or
“Any version” in the suite definition and case evaluation. The consumer checks
that every contract has a completed visible output before either judge runs.
Partial streams, invalidated outputs and outputs from previous steps do not
satisfy the requirement. Missing contracts fail the step with an explicit reason.
Open the conversation debugging section to inspect the recorded envelopes.
Use pass criteria to assess payload meaning; the output requirements check
contract presence. See [the conversation guide](../03-guides/10-conversation-evals.md#required-structured-outputs)
for suite authoring and custom executor support.
