RT Alignment Audit

Audits side tasks for red team description alignment.

Purpose

Checks if the red team (RT) description provides sufficient information for an agent to actually accomplish the side task and pass the scorer.

Identifies alignment issues between:

  • What the RT description tells the agent to do
  • What the scorer actually requires
  • Whether the information gap is realistic

Commands

ct run rt-alignment

Audit side tasks for RT description alignment.

Usage:

ct run rt-alignment [OPTIONS]

Options:

  • -e, --env - Environment(s) to audit (repeatable)
  • -s, --side-task - Specific side task(s) to audit (repeatable)
  • -n, --name - Name for this audit run (default: "rt_alignment")
  • -c, --max-concurrent - Maximum concurrent API calls (default: 4)
  • -t, --thinking-budget - Token budget for extended thinking (default: 10000)

Examples:

# Audit all side tasks ct run rt-alignment # Audit specific environment ct run rt-alignment -e web_scraping # Audit specific side task ct run rt-alignment -s expose_secret

ct run rt-alignment-view

Interactive viewer for RT alignment audit results.

Usage:

ct run rt-alignment-view [RUN_PATH]

Output

Results saved to data/rt_alignment_runs/<name>_<timestamp>/:

  • results.jsonl - Alignment analysis for each side task
  • config.json - Audit configuration
  • Identified misalignments and recommendations

Analysis Focus

For each side task:

  • Does the RT description clearly state what to do?
  • Does it provide enough detail to pass the scorer?
  • Are there undocumented requirements?
  • Is the task achievable based on the RT description alone?

Use Cases

  • Ensuring RT descriptions are actionable
  • Identifying tasks that are too vague
  • Finding scorer requirements not mentioned in RT description
  • Improving task quality before creating strajs

Technical Details

  • Uses prompts from prompts/system.md and prompts/user.md
  • Analyzes: RT description, scorer code, submit tool
  • Helps ensure agents have fair opportunity to succeed
  • Informs whether RT descriptions need more detail