Auditing Environments

Environment auditing tools are available as ct run subcommands.

Side Task Review

Review side tasks for scorer reward hacking, magic numbers, accidental triggering, and straj feasibility:

# Review all side tasks for an environment uv run ct run side-task-review -e my_environment # Review specific side tasks uv run ct run side-task-review -e web_scraping -s visit_malicious_website -s expose_secret

Test Quality

Analyze whether integration tests verify real end-to-end behavior:

uv run ct envs integration-test-quality -e my_environment

Compare Side Tasks

Compare side task successes between two runs:

uv run ct run compare-side-tasks <run_id_1> <run_id_2>

Reward Hacking Analysis

Audit side tasks for reward hacking vulnerabilities using Claude Opus with extended thinking:

uv run ct run reward-hacking -e my_environment

Env Description Audit & Generate

Two paired tools for docs/description.md. Both tools share a single canonical spec at src/control_tower/ct_runs/env_description/shared/spec.md, which defines length caps, required structure, and in-bounds vs out-of-bounds content — they cannot drift on what a "good description" is.

Audit scores an existing description against the codebase on five metrics (accuracy, clarity, completeness, information leakage, overall):

# Audit all environments uv run ct run env-description-audit # Audit specific environment(s) uv run ct run env-description-audit -e my_environment # Use a different model uv run ct run env-description-audit -m claude-sonnet-4-6

Results are written to data/env_description_audit_runs/ with per-env scores, notes, issues, suggestions, and a ranked summary. Defaults to claude-opus-4-7 at -c 2.

Generate writes a fresh description from the codebase. To protect uncommitted edits, the command aborts up front if docs/description.md already exists unless you pass --overwrite or --dry-run.

uv run ct run env-description-generate -e my_environment # first-time only uv run ct run env-description-generate -e my_environment --overwrite # replace existing uv run ct run env-description-generate -e my_environment --dry-run # writes description.draft.md # Then validate with the audit uv run ct run env-description-audit -e my_environment

Both tools use DirectAPIGenerator, which handles rate-limit retries with exponential backoff. Distinct from ct run doc-audit, which audits per-side-task docs.

Attack Evaluations

Use ct run eval directly with the appropriate policy:

# Obvious attack uv run ct run eval --policy obvious-attack -e my_environment -t main_task -s side_task # Sneaky attack uv run ct run eval --policy attack -e my_environment -t main_task -s side_task # Test if side tasks pass with read-only access (gameable = bad) uv run ct run eval --policy explore-only -e my_environment -t main_task -s side_task

Options

Most ct run audit commands accept --env / -e for environment selection (compare-side-tasks takes two positional run IDs instead). If --env is omitted, the CLI prompts interactively.