Goal Completeness Audit

Audits side tasks to verify agents have sufficient information to complete them.

Purpose

Checks whether the agent has sufficient information from the goal, scoring details, and submit tool description to successfully complete the side task and pass the scorer.

Compares what the scorer actually checks against what the agent is told to identify:

  • Missing information in goal descriptions
  • Undocumented scorer requirements
  • Ambiguous task specifications
  • Information asymmetries

Commands

ct run goal-completeness

Audit side tasks for goal completeness.

Usage:

ct run goal-completeness [OPTIONS]

Options:

  • -e, --env - Environment(s) to audit (repeatable)
  • -s, --side-task - Specific side task(s) to audit (repeatable)
  • -n, --name - Name for this audit run (default: "audit")
  • -c, --max-concurrent - Maximum concurrent API calls (default: 4)
  • -t, --thinking-budget - Token budget for extended thinking (default: 10000)

Examples:

# Audit all side tasks ct run goal-completeness # Audit specific environment ct run goal-completeness -e web_scraping # Audit specific side task ct run goal-completeness -s visit_malicious_website # Higher thinking budget for thorough analysis ct run goal-completeness -t 20000

ct run goal-completeness-view

Interactive viewer for goal completeness audit results.

Usage:

ct run goal-completeness-view [RUN_PATH]

Output

Results saved to data/goal_completeness_runs/<name>_<timestamp>/:

  • results.jsonl - Completeness analysis for each side task
  • config.json - Audit configuration
  • Identified gaps between scorer requirements and agent knowledge

Analysis Focus

For each side task, evaluates:

  • Is the goal description clear and actionable?
  • Does the submit tool description match scorer requirements?
  • Are there hidden scoring criteria not mentioned to the agent?
  • Can an agent realistically succeed based on provided information?

Technical Details

  • Uses prompts from prompts/system.md and prompts/user.md
  • Analyzes: goal text, submit tool definition, scorer code
  • Identifies information gaps that could cause legitimate failures
  • Helps improve task quality and fairness