HTTP API
The viewer is a single Next.js deployment that is both the website and the data plane. Every machine-boundary consumer — the ct CLI, browser client-components (lazy loads and downloads), and LLM agents — talks to the same /api routes documented here. The website's own server-rendered pages and server actions call the underlying data-layer functions in-process and do not go over HTTP, so they are not listed as routes.
This is the contract the CLI speaks and the surface a self-hosted viewer exposes. It is unversioned and lockstep: there is no /v1 prefix, breaking changes land atomically with their client, and a stale client fails loudly (a 400/404) rather than silently misbehaving. The remedy for a loud failure is to update your checkout (git pull).
Authentication
A deployment runs in exactly one mode, chosen by the VIEWER_MODE env var (private or public; any other value, including unset, is fail-closed and 500s every request).
- Private mode: every route requires a per-user token. Send it as an
Authorization: Bearer <token>header, or as thecontrol_tower_admin_sessioncookie that the web login sets. Tokens live in the deployment's own<MONGO_DB>_auth.userscollection; any valid token grants the full read/write surface (there are no scopes or roles yet). There is no HTTP Basic auth. - Public mode: reads need no credential; the data layer narrows what is visible. Writes are refused with
403by construction, and a few read routes that would expose full.evalarchives also 403 (see the download routes below).
Get a token with the CLI (ct login caches one locally; CONTROL_TOWER_API_TOKEN overrides the cache for CI), or directly:
curl -sX POST "$BASE_URL/api/auth/login" \ -H 'content-type: application/json' \ -d '{"username": "you", "password": "..."}' # => {"token": "..."}
A missing or invalid token on a private deployment returns 401 {"error": "Unauthorized"}. The CLI branches on status alone — 401 prints a "run ct login" hint.
Wire conventions
- snake_case everywhere on the wire. Query params, request-body fields, and response fields are all snake_case (
total_count,page_size,ids_only,task_set,run_id). camelCase is a TypeScript-internal convention that never appears in HTTP. - Unknown query params fail loudly. Each route declares its allowed param names; any param outside that set is a
400that names the offender and, when the fix is an obvious camelCase→snake_case slip, suggests the right spelling (e.g.{"error": "Unknown query param: 'pageSize' (did you mean 'page_size'?)"}). The suggestion is message text only, never an accepted alias. Routes that take no params reject any query param. - Uniform list shape. Every list response is
{"items": [...], "total_count": N, "page": P, "page_size": S, "total_pages": T}. - Every list response is bounded. No request returns a whole collection.
limit=Nis "the newest N" (page one at N);page+page_sizepaginate (both or neither); a request that supplies neither is page one at the maximum page size (500). All three clamp to that maximum.total_countis always the true match count andpage_size/total_pagesdescribe the window you got, so a caller that wants everything can see how much it did not get and page for the rest — it never has to know the server's ceiling. - Single documents are returned bare — the document itself, with no
{data, meta}envelope. - Errors are
{"error": "<human message>"}plus the correct HTTP status. There are no machine-readable error codes; clients branch on status (400,401,403,404,409,413,502).
Route surface
Trajectories
| Route | Purpose |
|---|---|
GET /api/trajectories | List trajectory summaries. Filter params come from the shared trajectory-filter vocabulary (environment, main_task, side_task, run_id, model, tag, tags_include, tags_exclude, main_task_outcome, side_task_outcome, dataset, trajectory_ids, versioned_dataset + versioned_dataset_version). main_task_outcome / side_task_outcome take pass, fail, or unscored (never scored); any other value is a 400. Paginate with page/page_size (together) or cap with limit. ids_only=true makes items a string[] of trajectory ids. |
GET /api/trajectories/{id} | The {metadata, sample} bundle: the Mongo metadata document plus the full eval sample brokered from S3. 404 if unknown, 502 if the sample can't be fetched. |
GET /api/trajectories/{id}/sample | The raw, still-compressed .eval sample bytes, streamed verbatim (see below). |
GET /api/trajectories/{id}/eval | 302 redirect to a short-TTL presigned S3 GET of the raw .eval archive (see below). |
GET /api/trajectories/zip | Streaming ZIP export. Requires a scope — run_id, dataset, or trajectory_ids — else 400. Optional sort_by/sort_dir, archive_name, and an inclusive 0-based zip_index_start/zip_index_end window. HEAD previews whether the request is valid/allowed. |
PATCH /api/trajectories/{id} | Merge-patch the editable fields (see PATCH semantics). Returns the updated {metadata} overlay only — a metadata edit never touches the sample, so it is not fetched or returned; use GET for the full {metadata, sample} bundle. |
POST /api/trajectories/register | Register a run header and/or trajectory documents (the CLI upload path; see register semantics). |
Runs
| Route | Purpose |
|---|---|
GET /api/runs | List runs, newest first. Filters: search, policy, task_set, uploader, producer, model, control_setting, tags_include, tags_exclude, dataset. Paginate with page/page_size, or cap with limit (providing both is a 400). |
GET /api/runs/{id} | One run document, bare. 404 if unknown. |
PATCH /api/runs/{id} | Merge-patch the editable fields (see PATCH semantics). Returns the updated run. |
Comments
Comments are their own append-only sub-resource, never a merge-patch field. The same Comment entity hangs off three parents — a run, a whole trajectory, and a single step (keyed by tool_call_id) — and all three expose the identical three methods.
| Route | Purpose |
|---|---|
POST /api/runs/{id}/comments | Append a comment to a run. |
PATCH /api/runs/{id}/comments/{comment_id} | Edit that comment's text. |
DELETE /api/runs/{id}/comments/{comment_id} | Soft-delete that comment. |
POST /api/trajectories/{id}/comments | Append a comment to a trajectory. |
PATCH /api/trajectories/{id}/comments/{comment_id} | Edit that comment's text. |
DELETE /api/trajectories/{id}/comments/{comment_id} | Soft-delete that comment. |
POST /api/trajectories/{id}/steps/{tool_call_id}/comments | Append a comment to one step of a trajectory. |
PATCH /api/trajectories/{id}/steps/{tool_call_id}/comments/{comment_id} | Edit that comment's text. |
DELETE /api/trajectories/{id}/steps/{tool_call_id}/comments/{comment_id} | Soft-delete that comment. |
Shared semantics across all nine:
- The client sends only
text.POSTandPATCHtake{"text": "..."}and nothing else; a body without a stringtextis a400.PATCHaccepts an empty string.DELETEtakes no body. - Identity and clock are server-owned. The comment's
id,created_by/created_at,edited_by/edited_at, anddeleted_by/deleted_atare all stamped from the authenticated actor and the server clock; a client cannot set or spoof them. - Delete is soft. The comment and its text are retained with
deleted_at/deleted_byset, so the audit trail survives. - Responses:
POSTreturns201with the new comment;PATCHreturns200with the updated comment;DELETEreturns200 {"ok": true}. An unknown parent, an unknown comment id, or a comment that is already soft-deleted is a404; atool_call_idoutside[A-Za-z0-9_-]+is a400.
Datasets
Datasets are named, versioned collections of run ids or trajectory ids. The version number is assigned by the server on publish (the client never picks it), which removes the read-then-insert race the old client-side path had.
| Route | Purpose |
|---|---|
GET /api/datasets | List datasets (latest-version summaries) with their lock state. Optional kind = trajectory or run. |
GET /api/datasets/{name} | The latest published version of a dataset. |
GET /api/datasets/{name}/versions | Every version, newest first. |
GET /api/datasets/{name}/versions/{v} | A single version (v is a positive integer). |
POST /api/datasets/{name}/versions | Publish a new version, in one of two body shapes. Absolute: {"item_ids": [...]} plus optional kind (trajectory or run, allowed only on this shape) and description — returns the new version document. Relative: {"add": [...]} and/or {"remove": [...]} plus optional description — the delta is merged server-side against the latest version (this is what ct dataset add / ct dataset remove send) and the response is {version, added, already_present, not_found, changed}. Supplying both shapes, or a delta that names no items, is a 400; so is any query param (the route accepts none). 409 if the dataset is locked. |
PUT /api/datasets/{name}/lock | Lock a dataset (optional {"reason": "..."} body), blocking membership writes into it. Returns the lock document. |
DELETE /api/datasets/{name}/lock | Unlock a dataset. Returns {name, was_locked}. |
Sabotage evals
| Route | Purpose |
|---|---|
GET /api/sabotage-evals | List sabotage evals, most recent first (graph images stripped for speed). |
GET /api/sabotage-evals/{id} | One sabotage-eval document including its base64 graph images (the browser lazy-loads this after the SSR shell paints). 404 on a malformed or unknown id. |
POST /api/sabotage-evals | Store a sabotage-eval result document exactly as produced by ct run sabotage-eval. The doc is schemaless by design — only its top-level shape (a JSON object) is checked — the server mints the ObjectId and returns 201 {"id": "<hex>"}. |
Eval logs (S3 brokering)
| Route | Purpose |
|---|---|
POST /api/eval-logs/upload | Plan an add-only upload for one .eval file: returns a single presigned PUT {url, bucket}. 409 if the object key already exists (uploads are add-only, never overwrite), 413 if size_bytes exceeds S3's 5 GB single-PUT limit. |
GET /api/eval-logs/head | Preflight: does filename already exist in S3? Returns {size_bytes, e_tag, version_id, bucket} or 404. |
Auth
| Route | Purpose |
|---|---|
POST /api/auth/login | {username, password} → {token}. Used by ct login; the web login page uses a server action, not this route. |
AI (Markdown for agents)
GET /api/ai/runs, GET /api/ai/trajectories, and GET /api/ai/trajectories/{id} return the same data as Markdown for LLM agents, over the same filters and auth as the JSON routes. /llms.txt is the agent-facing entry point: it self-describes the data model, endpoints, and wire contract, and agents re-discover it each session. It is the source of truth for the AI surface.
Raw sample byte stream
GET /api/trajectories/{id}/sample streams the raw, still-compressed .eval sample bytes verbatim over a chunked 200 — a format-agnostic byte proxy. The server does no decompression, parsing, or derivation: it forwards the compressed sample member straight from S3, so the ~1 MB compressed payload crosses the wire instead of the hundreds of MB it inflates to, and the function never holds the decompressed sample in memory. The client inflates and derives the transcript itself (a Web Worker, in the browser).
Response headers:
Content-Type: application/octet-stream.X-Sample-Compression: <codec>— names the codec so the client knows how to inflate (derived from the.evalZIP member's compression method).Content-Encoding: identity— the payload is already compressed; this stops the platform from re-encoding it.Cache-Control: no-store.
Failures split on when they happen relative to the committed 200:
- Before the stream starts (bad id, missing document, S3 open failure, or an unsupported ZIP compression method): a plain
{"error": ...}JSON body with a400/404/502status — no stream. - After the
200is committed: the status is already sent, so a failure can only truncate the byte stream. The client's derivation surfaces the truncation as an error — a transcript that decodes without a terminating summary is treated as truncated, not complete.
PATCH (merge-patch) semantics
PATCH /api/runs/{id} and PATCH /api/trajectories/{id} apply JSON merge-patch over a strict field whitelist, shared with the website's edit actions so the two can't drift:
- Runs:
name,tags,description,policy. - Trajectories:
description,tags,attack_analysis.
These are singleton fields only. Comments are their own sub-resource (see above) and dataset membership lives in a dataset version's item_ids (edited by publishing a version) — neither is patchable here.
Rules:
- An absent field is left unchanged.
- An explicit
nullclears a field, and is legal only on nullable fields (descriptionandpolicyon runs,attack_analysison trajectories); anullanywhere else is a400. A run'snamemay not be set empty. - Arrays replace wholesale — there are no add/remove operations, so read-modify-write is the client's job.
- Unknown fields are a
400, never silently ignored. - The response is the updated document.
PATCH /api/runs/{id}returns the run (same shape as itsGET);PATCH /api/trajectories/{id}returns the{metadata}overlay only — a metadata edit never touches the S3 sample, so it is neither fetched nor returned (useGETfor the full{metadata, sample}bundle, avoiding a post-write502if S3 is unavailable). An unknown id is a404.
Register semantics
POST /api/trajectories/register is the single write path for run headers and trajectory documents (the CLI's ct run eval, ct runs make, ct run monitor, and ct traj upload all end here). The body is {run?, trajectories, eval_log_sizes?, client, schema_version}.
- Schema handshake:
schema_versionmust equal the version the server supports, else400with a "your control-tower checkout is incompatible — rungit pull" message. This is the G4 fail-loud mechanism for the upload contract. - Eval logs must exist first: the server HEAD-verifies every referenced
.evalobject is already in S3 (add-only, S3-then-Mongo ordering); a missing object is a400. - Idempotent retries: if all posted trajectory ids already exist, the call returns
200 {"ok": true, "already_registered": true}— a lost response followed by a retry reads as the success it is. Partial overlap (some ids exist, some don't) stays a409because it is a genuine anomaly. - Registration carries no dataset membership, so there is no lock check here. Membership lives in a dataset version's
item_ids, and the lock gate sits on the publish path (POST /api/datasets/{name}/versions, a409when locked). Locks are an accident interlock, not authorization — there is no client-side check and no bypass flag; the only path into a locked dataset is to unlock it first. - The run header is create-once. Re-registering an existing
run_iddoes not rewrite the header — therun_idis the idempotency key, and header fields change only through the auditedPATCH. Trajectories are inserted before the run header is upserted, so a trajectory conflict never commits a run header pointing at missing trajectories. - The run header is optional;
run_idis a soft reference.runmay be omitted (ct traj uploadposts trajectories withrun=null), in which case only the trajectory documents are written and no run header is created. Each trajectory keeps its client-suppliedrun_idverbatim — it is never force-emptied and never triggers an auto-created run header — so arun_idthat resolves to no run is a legitimate shape. Readers treat the reference as soft: run-less trajectories list and render normally, and a danglingrun_idnever fabricates a run (GET /api/runs/{id}for such an id is a404, and the trajectory-detail "Run" link points at a plain/runs/<run_id>404).
Presigned .eval downloads
Raw .eval archive reads are brokered, never streamed through the function: GET /api/trajectories/{id}/eval authenticates the token, then returns a 302 redirect to a short-TTL presigned S3 GET (Cache-Control: no-store). The client — browser or CLI — follows the redirect and downloads straight from S3, so no AWS credentials ever leave the server and the API token stays the only long-lived client credential. The CLI follows the redirect manually so its Authorization header never travels to S3.
In public mode this route returns 403 ("Full .eval archive download is unavailable in public mode"), as does the run_id-scoped ZIP export — public deployments do not hand out whole eval archives.
Operator workflow (ops/)
Nothing in the application or the CLI provisions Mongo indexes, seeds data, or manages login accounts on its own. Those are deliberate, out-of-band operator steps performed by someone holding write credentials for a cluster, and they live in the standalone ops/ package (the single ct-ops console script, with a subcommand per step): index provisioning, viewer login accounts, cluster copy for the viewer cutover, and local seed data. ops/ is a separate package with its own venv and zero control_tower imports. See ops/README.md for the full workflow.
Self-hosting the viewer
There is no separate "API client" viewer mode and no second implementation. A downstream user who wants their own viewer self-hosts the same Next.js app over their own MongoDB — their deployment is their API, and their CLI points at their deployment's URL (CONTROL_TOWER_API_BASE_URL). MONGOURI lives only in each deployment's own server environment (Vercel env, a self-hosted container, or ops/); it is never handed to a client. That MongoDB must be a replica set (single-node is fine), not a standalone mongod, so register's transaction has somewhere to run — see Database — replica-set requirement. See Web Viewer for the self-hosting quickstart and packages/viewer/README.md for the library shape and env vars.