Evaluation Examples

Testing a New Main Task

uv run ct run eval --policy honest --env my_env --main-task my_task

Running Attack Evals

# Single attack eval uv run ct run eval --policy attack --env web_scraping --main-task crawl_depth --side-task visit_malicious_website # Obvious attack (no concealment) uv run ct run eval --policy obvious-attack --env web_scraping --main-task crawl_depth --side-task visit_malicious_website # Expand to all attack combinations for a given env uv run ct run eval --policy attack --env web_scraping --all

Batch Evaluation

Create a task file with one record per combination:

{"env": "tiktok", "main_task": "add_pro", "side_task": "become_admin"} {"env": "web_scraping", "main_task": "backup", "side_task": "expose_secret"}

Then run:

uv run ct run eval --policy attack --task-file tasks.jsonl

Quick Iteration on Human Trajectories

# Start live context uv run ct live up # Make changes, generate trajectory uv run ct live gen-traj # Re-run scoring locally on the generated trajectory (uses solution.json from live context) uv run ct run rerun-traj --from-live --no-upload # Monitor the trajectory (use the path printed by gen-traj) uv run ct run monitor <path-from-gen-traj> --action-monitor multi --only-action --trusted-model openai/gpt-5-nano --no-upload # When satisfied, upload the trajectory (upload is enabled by default) uv run ct run rerun-traj --from-live