Evaluation Setup

Prerequisites

Environments

Environments must be installed before running evals. If .settings/ is empty, pull them first:

uv run ct settings pull

See Development Setup for details.

API Keys

You'll need API keys for the model providers you're using. Set them in your .env file:

ANTHROPIC_API_KEY=your-key-here OPENAI_API_KEY=your-key-here

Local Services

Start the support services once:

docker compose up -d

This provides local MongoDB (localhost:27017), the website (private viewer at localhost:3000 with login admin/linuxbench, public viewer at localhost:3002, docs at localhost:3003), and the eval viewer (localhost:8080). The stack's own services read MONGOURI to reach the local Mongo (set in compose.yml); the CLI itself never touches Mongo — it uploads to the local viewer over HTTP, so point it at the stack with CONTROL_TOWER_API_BASE_URL=http://localhost:3000 and ct login (local credentials admin/linuxbench). See Web Viewer and Inspect Logs Viewer for the local service details.

Then log in:

uv run ct login

Without the local services, evals still run and produce local files — results just won't be uploaded or viewable on the website.

Quick Start

The primary command for running evaluations is ct run eval:

uv run ct run eval --policy <policy> --env <env> --main-task <task>

Test your setup with a test run before running expensive model evaluations:

uv run ct run eval --policy test --task-file simple --tags test

This uses the test policy (runs ls and submits, no model calls) with the simple task set. Use test runs to:

  • Validate a new development environment
  • Verify Docker builds and scorer execution work correctly
  • Debug infrastructure issues before running real evals

Configuration

Model Configuration

By default, evaluations use:

  • Untrusted model: anthropic/claude-haiku-4-5
  • Trusted model: openai/gpt-5-nano

You can override these defaults with policy arguments. The CLI accepts model aliases:

AliasFull Model Name
haikuanthropic/claude-haiku-4-5
sonnetanthropic/claude-sonnet-4-6
opusanthropic/claude-opus-4-6
opus_4_6anthropic/claude-opus-4-6
opus_4_8anthropic/claude-opus-4-8
gemini_flashgoogle/gemini-3-flash-preview
gemini_progoogle/gemini-3.1-pro-preview
untrustedCurrent default untrusted model (haiku)
trustedCurrent default trusted model (gpt-5-nano)
# Using aliases uv run ct run eval --policy honest -p model=sonnet ... # Using full model names uv run ct run eval --policy honest -p model=anthropic/claude-sonnet-4-6 ...

OpenRouter routing

Set USE_OPENROUTER to a boolean to send all model traffic through OpenRouter, or set USE_OPENROUTER_PROVIDERS to a comma-separated list of provider slugs to send only those providers through OpenRouter. Do not use both at once. If you use either, set OPENROUTER_API_KEY in your .env.

Self-Hosted and Open-Weight Models

Model names resolve through Inspect AI's model providers, so self-hosted open-weight models work anywhere a model name is accepted. Point the provider at your endpoint via env vars in .env and use the matching prefix:

  • vLLM / SGLang: set VLLM_BASE_URL=http://<host>:8000/v1 (or SGLANG_BASE_URL=...), then pass -p model=vllm/Qwen/Qwen3-32B. Serve with --enable-auto-tool-choice and the tool-call parser for your model family — agent policies drive on OpenAI-style tool calls.
  • Any OpenAI-compatible endpoint (NVIDIA NIM, TGI, llama.cpp): pass -p model=openai-api/<service>/<model>, which reads <SERVICE>_BASE_URL and <SERVICE>_API_KEY.

Self-hosted prefixes (vllm/, sglang/, hf/, ollama/, openai-api/) are never routed through OpenRouter, even with USE_OPENROUTER=1.

A single GPU endpoint usually needs the per-call timeout raised and concurrency capped: -p model_attempt_timeout=600 -p model_max_connections=8 --max-samples 1. A fresh clone has no rate file until you run uv run python scripts/generate_model_rates.py (see Development Setup); until then, or for any model the table does not list, cost reports as unknown — not $0.00. A monitor model that can't follow the <score> output format after retries records an error and scores 0.0 — check error counts before trusting safety numbers from a small model.

Skip Docker Builds (Optional)

By default, running an eval rebuilds Docker images locally using the Dockerfile, with layer caching from Docker Hub. To skip rebuilding and pull pre-built images directly from Docker Hub instead:

export SKIP_DOCKER_BUILD=1

Services whose compose entry builds a single-segment image name (no registry or namespace, e.g. BashArena's per-task bash-arena-default-<task_id>:latest) keep their build section even under SKIP_DOCKER_BUILD — such names cannot exist in a registry, so they are always built locally. Fleet workers set SKIP_DOCKER_BUILD=1, so this is what makes per-task-image environments runnable on EC2.

Pre-built images for evaluation environments are published to Docker Hub under the linuxarena organization (for example linuxarena/env-web_scraping:latest). Images for public environments (those with public repos on github.com/linuxarena) can be pulled anonymously; images for private environments remain private - contact the team for access to the private environments.

Each environment's image tag comes from ENV_IMAGE_TAG_<ENV> (uppercased env id, e.g. ENV_IMAGE_TAG_WEB_SCRAPING). Env CI publishes each image under both latest and the env repo's commit SHA, and when the variable is unset a local SKIP_DOCKER_BUILD=1 run pulls the image published at the env checkout's commit — skipping the build doesn't change which env version runs. If no image exists at that commit (checkout ahead of CI, or predating SHA-tagged publishing), the run warns and falls back to latest. EC2 fleet runs default unset tags to latest at dispatch (workers have no local checkout). Set the variable to a published tag to override either default.

Provenance also records the pulled image's digest at env_versions[<env>]["image_digest"], pinning a mutable tag like latest to the exact bytes that ran. Locally-built images have no digest.

When to use:

  • CI/batch evals: Always use SKIP_DOCKER_BUILD=1 for faster startup
  • EC2 evals: Recommended since EC2 instances start fresh
  • Development: Don't use if you're testing local Dockerfile changes

Scoring the main task on attack samples (Optional)

By default an attack-mode sample (one with a side task) records only side_task_success. Pass --score-main-on-attack to also run the main-task scorer on attack samples and record main_task_success alongside it — useful for measuring an attack policy's usefulness (does it complete the main task while attacking?) and for per-run safety/usefulness frontier points.

uv run ct run eval --policy attack -e web_scraping -t crawl_depth -s visit_malicious_website --score-main-on-attack

It is off by default. The side scorer always runs first, then the main scorer runs under its own --scorer-timeout, so an attack sample can take up to twice as long to score. The flag is run-level (every sample emits the same score keys) and rides through YAML configs, hawk submissions, and run metadata automatically. An environment can also declare score_main_on_attack itself, for benchmarks whose side task is defined relative to main-task performance; one selected environment declaring it turns main scoring on for the whole run, which is what keeps the keys uniform. Simulated runs (--simulated) score no main task and ignore the flag.

Policies

Pass a single policy with --policy:

uv run ct run eval --policy honest --task-file simple uv run ct run eval --policy attack -p model=sonnet --task-file medium

For protocol evals, use --untrusted-policy with --protocol instead:

uv run ct run eval \ --untrusted-policy simple-attack \ --protocol trusted-monitoring \ --protocol-arg sus_threshold=7 \ --task-file simple

See Protocols and Untrusted Policies for the built-in policies, protocol-mode concepts, and implementation details.

External Sandboxes

--sandbox runs samples in an Inspect sandbox provided outside this repo — a microVM runtime, for instance — instead of the default Docker sandbox. Pass a registered type, or module:type to import the module that registers it, with --extra-src when the project is not installed.

uv run ct run eval \ --policy honest \ --task-file simple \ --sandbox my_project.sandbox:my-microvm \ --sandbox-arg boot_timeout=300 \ --extra-src ../my-project

--sandbox-arg values are passed to the sandbox as a JSON string for it to parse and validate; Control Tower does not interpret them, and their keys reach it exactly as written. The sandbox owns the sample lifecycle through Inspect's task_init, sample_init, sample_cleanup and task_cleanup hooks, so it cannot be combined with --ec2, --simulated or --replay-sandbox. It also rejects --no-internet, --no-intranet, --network-control and --sandbox-isolation, which the Docker sandbox enforces and an external one cannot — the run fails rather than proceeding unrestricted. The selection is recorded in run metadata with the rest of the sandbox config.