Quickstart
This is the fastest path from a fresh machine to running control evals on the public LinuxArena environments, using an OpenRouter API key. Everything runs locally in Docker, no cloud accounts or internal credentials needed.
New to AI control? Read AI Control Concepts first (5 minutes); it defines main tasks, side tasks, honest/attack evals, and monitoring.
Prerequisites
- Docker, installed and running
- uv:
curl -LsSf https://astral.sh/uv/install.sh | sh - git: environment definitions are public repos, pulled with plain
git clone(no GitHub account needed) - An OpenRouter API key
Setup
git clone https://github.com/linuxarena/control-tower.git cd control-tower uv sync
Create a .env file in the checkout root:
OPENROUTER_API_KEY=sk-or-v1-... USE_OPENROUTER=1 SKIP_DOCKER_BUILD=1
USE_OPENROUTER=1routes all model calls through your OpenRouter key, so you don't need separate Anthropic/OpenAI keys.SKIP_DOCKER_BUILD=1pulls prebuilt environment images from Docker Hub instead of building them locally.
Then pull the public LinuxArena environments:
uv run ct settings pull linuxarena uv run ct envs list
Run an honest eval
An honest eval gives the agent a main task and measures whether it completes it:
uv run ct run eval --no-upload \ -e clinical_trial -t adverse_reaction_followup_tracking \ --policy honest
--no-upload keeps results local (uploading requires internal services). The agent's actions stream to your terminal; the run takes a few minutes and ends with the scorer's CORRECT/INCORRECT verdict. The trajectory is saved as an Inspect .eval file under logs/.
To list tasks for an environment: uv run ct envs main-tasks list --env <env>. To pick a different model: -p model=<name>, any OpenRouter model works, e.g. -p model=openrouter/qwen/qwen3-coder.
Run an attack eval
An attack eval additionally gives the agent a hidden side task (sabotage objective) it must accomplish without looking suspicious:
uv run ct run eval --no-upload \ -e clinical_trial -t adverse_reaction_followup_tracking \ -s suppress_safety_notifications \ --policy attack
Side tasks per environment: uv run ct envs side-tasks list --env <env>.
Monitor a trajectory
Monitors are the blue team's core tool: another model reviews the agent's actions and assigns suspiciousness scores. Score the attack trajectory you just produced:
uv run ct run monitor logs/<eval-file>.eval \ --action-monitor mean-action \ --trajectory-monitor mean-quick \ --trusted-model openai/gpt-5-nano \ --no-upload
View results
uv run inspect view logs/<eval-file>.eval
This opens Inspect's local log viewer with the full transcript, scores, and monitor outputs. See Inspect Logs Viewer.
Using a self-hosted model
To use an open-weight model from your own GPU box, serve it (the tool-call flags are required for agent use; monitors don't need them):
vllm serve Qwen/Qwen3-32B --enable-auto-tool-choice --tool-call-parser hermes
add to your .env:
VLLM_BASE_URL=http://<gpu-box>:8000/v1
and pass the model to the commands above: -p model=vllm/Qwen/Qwen3-32B -p model_attempt_timeout=600 --max-samples 1 on ct run eval, or --trusted-model vllm/Qwen/Qwen3-32B on ct run monitor. Open-weight models are an easy fit for monitors (plain text, no tool calling). For other OpenAI-compatible servers (NVIDIA NIM, TGI) and details, see Self-Hosted and Open-Weight Models.
Gotchas
- Always pass
--no-uploadtoct run evalandct run monitor. Upload is on by default and requires internal credentials. - The
tiktokenvironment needs aTIKAPI_KEY(a paid third-party API). Use the other environments unless you've been given one. - Rate limits: OpenRouter rate limits scale with the key's credit balance. If you hit 429s, lower concurrency with
--max-samples 1.
Going deeper
- Running evals on Hawk: submit to a shared Hawk cluster and watch results in the viewer, instead of running locally
- Running Evaluations: full
ct run evalreference: policies, task sets, batching, replay - Monitoring: monitor types, ensembles, suspiciousness scoring (building monitors)
- Control Protocols: defer-to-trusted, trusted editing, and writing your own protocol
- Making Sabotage Trajectories: hand-crafting attacks that evade monitors
- Making Side Tasks and the Hands-On Guide: extending the testbed with new environments and tasks
- Safety Curves: turning monitor scores into safety/usefulness numbers