Run your first workflow eval
Updated 2 days ago • October 5, 2026
Start here if you already have a working Kortyx agent and want to run a suite from Studio. This recipe uses your existing agent and its normal model/tool configuration. You will create two application files, mount one endpoint, and configure the Studio API. No suite publication job is required.
For a standalone starting point, run the catalog endpoint example. It contains a real model/tool workflow, suite, HTTP server and optional local judge. The fixture verifies wiring; use your own permission path for application acceptance.
The execution path is Studio → Studio API → your application's eval endpoint → your existing agent → judge → saved result. Deploying Studio alone does not mount or enable the application endpoint.
1. Define one case
Create src/evals/suites.ts in your application:
Choose a product your test account can actually read. Author the scenario from known data; do not discover random products during startup. A suite is JSON data, not authentication code. Add more cases to this array as you learn about failures.
For this criterion, enable evidence on the workflow's existing useReason call:
Without emitted tool results the judge cannot establish that the reported price came from the lookup. This flag captures evidence; normal agent telemetry is configured separately if you want workflow costs and workflow inspection links.
2. Create the eval instance beside your agent
Create src/evals/index.ts. Export the eval instance without starting a server:
This minimal instance is sufficient when your workflow needs no caller-specific
context. If your tools depend on an authenticated user, add setup and execute
before running it:
authenticateTestActor and withApplicationRequest above are your application's
existing helpers, not Kortyx exports. The first obtains a test identity; the
second binds the same permission checks, request-scoped tools and context that
ordinary requests use. Adapt their arguments to your app. Reuse that existing
logic rather than implement a second permission system for evals.
Configure the test account in the application service. It may use a username and password, a refreshable token, or another identity mechanism your app already supports. Studio neither obtains that token nor assigns the account's roles. The eval service key authorizes initiation; it does not grant domain data access. Never put user tokens or passwords in suite definitions, progress or judge inputs.
Add teardown only if setup acquires resources that need releasing. It is not necessarily a sign-out operation. See application authentication and execution for cancellation, typed callbacks and interrupt responders.
3. Mount and deploy the application endpoint
For Next.js App Router, create src/app/api/evals/route.ts:
Its URL is https://your-app.example/api/evals. Keep a single handler instance
so its per-process active-run limit is shared across requests.
For a Hono Node server, mount the same Web handler before starting the server:
For Express/Fastify or another Node server, use the framework's Web Request/ Response bridge. It must preserve the request abort signal, stream the returned body, and cancel its reader when the client disconnects. Do not buffer the full POST response or redirect it through the interactive chat route. See the runnable Node adapter for a complete application using Hono's Node server bridge.
Generate a separate service key, store it privately, and supply it to the app:
The HTTP route explicitly reads EVAL_SERVICE_KEY; the local eval module does not need it. Kortyx does not automatically
read application env variables or enable routes. If your application has its
own eval feature flag, enable it in the deployed application. Ensure the server
and reverse proxy expose the exact GET/POST path in the target configuration.
Restart/deploy the app with its service key, test-account configuration and normal workflow model/tool configuration. A healthy container does not prove that the route was registered.
4. Register the application on Studio
On the Studio API backend, mount a private JSON file:
Use the existing Studio organization and project IDs. The environment must be allowed in that project; use its real configured label rather than infer one from the hostname. The application filter gets its ID/name from this file.
For Compose, apply this override alongside your installation's normal stack:
Keep the target file and env files out of Git and restrict their filesystem
permissions. Rerun the existing bootstrap job with the same stored Studio key
to grant eval:run, preserving its key and pepper. Run the release's normal
database migrations, then restart the API to load target changes. Hosted systems
use their normal secret mounts and bootstrap process; a host .env value alone
does not create a container mount.
For an app running on your host with Studio in Docker Desktop, use
http://host.docker.internal:YOUR-APP-PORT/api/evals and
"allowInsecureHttp": true on that target. Remote deployments use HTTPS.
localhost inside the API container points to that container.
Studio fetches the registered suites through authenticated GET discovery.
It does not receive suites through kortyx topology push. Changing code-authored
suites requires deploying the application and refreshing discovery.
5. Configure one judge
This recipe selects Studio judge. Configure these on the Studio API and restart it:
Choose a model/provider route supporting structured JSON verdicts. The browser
and consumer do not receive this key. This model grades execution; it does not
replace your workflow model. You can instead configure an App judge with
createEvalJudge and explicitly select it in Studio or with --judge app.
See judge selection.
Configuration belongs to different services
The CLI uses a Studio project key, not the application service key or an LLM provider key. Preserve the existing database URL and API-key pepper. The configuration reference lists all defaults, optional judge ID/version overrides and existing Studio authentication settings.
6. Diagnose setup before executing
Use the CLI from the release containing the doctor command. Connect it to the Studio API URL, not the browser's URL:
For a configured connection, use --connection staging instead of the direct
URL/key flags. The connection's default environment applies unless overridden.
Example when the consumer route is missing:
Doctor checks project read/execution access, matching targets, environment access,
authenticated manifest validity, registered suites and advertised judge support.
It starts no workflows, judge calls or saved eval runs. The consumer's GET
wrapper may perform its own authentication; the SDK's manifest handler does not
call setup. Provider credentials, test-user tool permissions and worker health
still need a real run. An older Studio API may return only “unavailable”; doctor
reports that limitation rather than invent a cause.
--judge app checks the code judge instead; --json returns one report with
schemaVersion, status and checks. Each check has id, status, message
and an optional remedy. Exit 0 means all configuration checks passed; exit
1 means a failure. Failed prerequisites produce skipped downstream checks.
Doctor uses discovery GET only; do not use it as a provider-token validation test.
7. Run and verify the actual workflow
Open Evals → Suites, refresh, and choose Catalog lookup. Select Studio judge and run the one case once. Or enqueue through the same configured API:
The command returns a run ID immediately. Inspect it in Studio, or with
kortyx studio evals runs get <run-id> using the same connection flags.
An enqueue response confirms a queued record, not a completed eval.
Before calling the integration complete, verify:
- The real deployed suite is visible, not merely an application connection label.
- The test actor reaches the normal tools with its intended roles and data access.
- The workflow runs, the selected judge grades it, and the final reasons/evidence are readable.
- Reloading the run retains its definition, observations and final result.
- Live progress arrives; workflow/judge costs appear when their usage is recorded.
- Cancelling a suitable test run reaches a saved terminal state and honors cleanup.
A behavioral failure can be a useful successful integration test: inspect the reason rather than weaken criteria to obtain a green result. Authentication, provider or execution errors need resolving before assessing behavior. Use a representative read-only case for onboarding; a passing fixture alone does not prove your real application's permissions or domain logic.
Once this works, add interrupt scenarios and required structured outputs, local terminal execution, or optional post-deployment CI triggering. Keep CI triggering after application deployment/readiness; it need not block releases.