ct run find-questions
Discover answer-question main tasks in an environment.
Usage
ct run find-questions [OPTIONS]
Options
| Option | Description |
|---|---|
-e, --env, --environment TEXT | Environment(s) to discover questions in (repeat for multiple) [required] |
--ec2 | Run evaluation on EC2 instead of locally |
--ec2-instance-type TEXT | EC2 instance type (only used with --ec2) [default: c7i.4xlarge] |
--ec2-region TEXT | AWS region for EC2 (only used with --ec2) [default: us-east-1] |
--keep-ec2-instance | Keep EC2 instance running after eval (for debugging) |
--startup-retries INTEGER | Number of times to retry Docker sandbox startup on transient failures [default: 7] |
--max-workers INTEGER | Maximum number of EC2 worker instances (--ec2 fleet mode) [default: 64] |
--concurrency-per-worker INTEGER RANGE | Max concurrent samples per worker (--ec2 fleet mode); workers start at 10 and adapt within [1, max] from observed rate-limit pressure [default: 32; x>=1] |
--estimate | Print cost estimate and exit without running (--ec2 fleet mode) |
--yes | Skip 5-second confirmation wait (--ec2 fleet mode) |
--new-fleet | Force fresh fleet run, skip auto-resume (--ec2 fleet mode) |
--no-internet | Block outbound internet and internet simulator access. |
--no-intranet | Block access to compose sibling services (e.g. Redis). |
--network-control | Grant NET_ADMIN to the agent container without applying restrictions at startup. Protocols that manage iptables mid-trajectory (e.g. privilege-monitor) need this. |
--sandbox-isolation | Wrap every agent command in a bubblewrap sandbox: read-only root, writable workspace only. A static bwrap is fetched automatically; set CT_BWRAP_HOST_BINARY to override with your own. |
--spot | Use EC2 spot instances for fleet workers (~60-90%% cheaper, with automatic on-demand fallback) |
--extra-src DIRECTORY | Directory of experiment code added to PYTHONPATH locally, so external --policy my.module:fn and --sandbox my.module:type refs resolve without committing into the repo; fleet mode (--ec2) also ships it to workers. Repeatable; each dir must have a distinct basename. |
--worker-setup TEXT | Script each fleet worker runs as root before taking jobs, named as <extra-src basename>/<path> inside an --extra-src tree (--ec2 only). Installs the host side of a --sandbox, for example. |
--sandbox TEXT | Run samples in an Inspect sandbox provided outside this repo, as a registered type (my-sandbox) or module:type to import the module that registers it. Replaces the default Docker sandbox; with --ec2 each worker runs it for its own jobs. |
--sandbox-arg TEXT | Config for --sandbox in key=value format (multiple allowed), passed through to the sandbox as JSON. |
--simulated | Run against an LLM-simulated sandbox instead of provisioning the real environment. Every tool output is fabricated by a simulator agent. |
--simulator-model TEXT | Model for the simulator agent (--simulated). [default: anthropic/claude-opus-4-8] |
--simulator-judge-model TEXT | Model for the transcript judge that scores simulated samples (--simulated). [default: anthropic/claude-opus-4-6] |
--replay-sandbox | Serve each action's recorded tool output from the replayed trajectory instead of executing it (requires --trajectory-id). Nothing is provisioned or executed. Works with any replay-like policy — one whose executed tool calls stay a verbatim prefix of the recording; a call that diverges raises. |
--simulator-scratch-mode [filesystem|default_image|full|none] | Which real container the simulator investigates in (--simulated): filesystem (codebase copy, free), default_image (env's default image only), full (the whole compose stack), or none (no container — no-code-execution ablation). [default: filesystem] |
--epochs INTEGER | Number of epochs to run |
--max-samples INTEGER | Maximum samples to run in parallel |
--log-dir TEXT | Directory for eval logs |
--retry-on-error INTEGER | Number of times to retry on error |
--no-fail-on-error | Continue running if individual samples error (instead of failing the whole run) |
--scorer-timeout INTEGER | Timeout in seconds for scorer execution |
--diagnostics TEXT | Comma-separated diagnostic scorers to run (default: none) |
--score-main-on-attack | Also run the main-task scorer on attack-mode samples (adds main_task_success). Off by default; scores each attack sample twice, so worst-case scoring time doubles. |
--token-limit INTEGER | Token limit for the evaluation |
--message-limit INTEGER | Message limit for the evaluation |
--sample-time-limit INTEGER | Time limit in seconds for each sample |
--dry-run | Show discovered questions without creating task directories |
--no-upload | Disable uploading eval results |
-p, --policy-arg TEXT | Policy args in key=value format |
--help | Show this message and exit. |