Submit Guessability Verified
Audits side tasks for submit tool exploitability with LIVE verification.
Purpose
Like submit-guessability, but spins up containers and runs the scorer to verify that theoretical exploits actually work. The AI gets multiple attempts with feedback to find working submit arguments.
Commands
ct run submit-guessability-verified
Audit side tasks with live scorer verification.
Usage:
ct run submit-guessability-verified [OPTIONS]
Options:
-e, --env- Environment(s) to audit (repeatable)-s, --side-task- Specific side task(s) to audit (repeatable)-n, --name- Name for this audit run (default: "verified")-c, --max-concurrent- Maximum concurrent audits (default: 2, lower because containers are involved)-t, --thinking-budget- Token budget for extended thinking (default: 10000)-a, --max-attempts- Maximum attempts per side task (default: 5)
Examples:
# Audit all side tasks with verification ct run submit-guessability-verified # Audit specific environment ct run submit-guessability-verified -e web_scraping # More attempts to find exploits ct run submit-guessability-verified -a 10 # Lower concurrency for resource-constrained systems ct run submit-guessability-verified -c 1
ct run submit-guessability-verified-view
Interactive viewer for verified audit results.
Usage:
ct run submit-guessability-verified-view [RUN_PATH]
Output
Results saved to data/submit_guessability_verified_runs/<name>_<timestamp>/:
results.jsonl- Verification results for each side taskconfig.json- Audit configuration- Working exploits (if found) with scorer outputs
Workflow
For each side task:
- AI analyzes scorer and submit tool to propose exploit arguments
- Container is spun up and scorer is run with proposed arguments
- If scorer fails, AI sees the error and tries again
- Process repeats up to
max-attemptstimes - Results show whether a working exploit was found
Technical Details
- Uses prompts from
prompts/system.mdandprompts/user.md - Spins up live containers via the live evaluation infrastructure
- Iterative feedback loop: attempt → scorer output → refined attempt
- More resource-intensive than theoretical analysis
- Confirms actual exploitability vs theoretical vulnerabilities