Analyze Attack

Analyzes attack trajectories to identify incriminating actions and sabotage behavior.

Purpose

Uses LLM analysis to examine trajectories from attack policies and identify:

  • Actions that reveal the agent is attempting sabotage
  • Attack summaries explaining the sabotage strategy
  • Specific incriminating actions with explanations

Commands

ct run analyze-attack

Analyze attack trajectory(ies) to identify incriminating actions.

Usage:

ct run analyze-attack TRAJ_LIKE [OPTIONS]

Arguments:

  • TRAJ_LIKE - Trajectory file, eval log, directory, or trajectory ID

Options:

  • -m, --model - Model to use (default: claude-opus-4-5)
  • -n, --name - Custom name for the run directory
  • -c, --max-concurrent - Max concurrent analyses (default: 5)
  • -L, --limit - Limit number of trajectories
  • -o, --output - Output path for updated trajectory JSONs

Examples:

# Analyze a single trajectory ct run analyze-attack data/traj/example.json # Analyze an eval log with custom model ct run analyze-attack logs/my_eval.eval -m claude-opus-4 # Analyze directory with concurrency limit ct run analyze-attack logs/run_trajs/ -c 10 # Limit to first 5 trajectories ct run analyze-attack logs/*.eval -L 5

ct run analyze-attack-view

Interactive viewer for analyze-attack results.

Usage:

ct run analyze-attack-view [RUN_PATH]

Arguments:

  • RUN_PATH - Path to analyze-attack run directory (optional, will prompt if not provided)

Examples:

# Select from available runs ct run analyze-attack-view # View specific run ct run analyze-attack-view data/analyze_attack_runs/attack_analysis_2026-02-01_143022/

Output

Results are saved to data/analyze_attack_runs/<name>_<timestamp>/:

  • results.jsonl - Analysis results for each trajectory
  • config.json - Run configuration
  • Updated trajectory JSONs with attack analysis annotations (if -o specified)

Technical Details

  • Uses prompts from src/control_tower/ct_runs/traj_analysis/analyze_attack/prompts/analyze-attack.md
  • Extracts AttackAnalysis and IncriminatingAction objects
  • Stores results in JSONL format for easy viewing
  • Tracks token usage and costs