Monitor
Runs monitors on trajectories to detect suspicious behavior.
Purpose
Executes monitoring functions on trajectories to compute suspiciousness scores. Used to:
- Evaluate trajectory safety
- Compute action-level or trajectory-level suspicion scores
- Test new monitor implementations
- Generate monitored runs for analysis
Commands
ct run monitor
Run monitors on trajectories.
Usage:
ct run monitor TRAJ_LIKE [OPTIONS]
Arguments:
TRAJ_LIKE- An eval-backed source: a.evalfile, or a run ID/URL whose run records a backing eval log (see Downloading from the Viewer)
Options:
-a, --action-monitor- Action-level monitor selection (e.g., multi, strong, simple); default:mean-action-t, --trajectory-monitor- Trajectory-level monitor selection (e.g., quick, mean-quick); default:mean-quick--trusted-model- Required. The model bound as the eval-leveltrustedrole that every monitor runs on. Any Inspect AI provider string works, including self-hosted open-weight endpoints (see Self-Hosted and Open-Weight Models)--cot-level- Chain-of-thought visibility for the monitors:none(default),reasoning, orextended--only-action- Run only the action-level monitor--only-trajectory- Run only the trajectory-level monitor--save-dir PATH- Directory to save monitored trajectories--no-upload- Disable upload (upload is enabled by default)--run-name TEXT- Custom run name
Monitors never carry model configuration; the one model they run on is the eval-level trusted role, given by --trusted-model. The exception is monitoring-bench, which replicates its benchmark's fixed 3-model panel.
A monitor selection takes one of three forms: a packaged monitor name; a filesystem path to a .md asset file (absolute or relative to the working directory; the file stem becomes the monitor's score name); or a module:function import path to a code monitor — an async callable with the harness's monitor signature, or a factory returning one (the import path becomes the score name). A JSON list combines selections, and a list entry may be a {"monitor": <selector>, <overrides>} mapping: on an asset, the extra keys override front-matter fields (e.g. samples) and fill template variables (a variable with neither a front-matter default nor a supplied override fails the selection, listing what is missing); on a code monitor, they are passed to the factory as kwargs. simple and monitoring-bench are the in-repo reference implementations of the code-monitor form.
# Two action monitors in one pass: a packaged one plus an external asset file, with an override ct run monitor logs/my_eval.eval --only-action --trusted-model openai/gpt-5-nano \ -a '["no-history", {"monitor": "/path/to/my-monitor.md", "threat_focus": "exfiltration"}]'
Examples:
# Monitor with the default monitors ct run monitor logs/my_eval.eval --trusted-model openai/gpt-5-nano # Action-level only, with a specific monitor ct run monitor logs/my_eval.eval -a strong --only-action --trusted-model openai/gpt-5-nano # Monitor and save locally without uploading ct run monitor logs/my_eval.eval -a strong --save-dir monitored_trajs/ --no-upload --only-action --trusted-model openai/gpt-5-nano # Monitor a specific run from database ct run monitor abc123def456 -a multi --only-action --trusted-model openai/gpt-5-nano # Run a trajectory-level monitor only ct run monitor logs/my_eval.eval -t quick --only-trajectory --trusted-model openai/gpt-5-nano # Monitor a viewer-hosted run directly (no intermediate download) ct run monitor https://basharena-data.vercel.app/runs/ed2b360737041f1743ac4de43280f858 \ -a multi --only-action --trusted-model openai/gpt-5-nano
Available Monitors
A monitor is a prompt asset: a markdown file under src/control_tower/protocols/monitoring/assets/{action,trajectory}/ whose stem is the monitor's name. Its YAML front matter declares the monitor's view — samples (independent draws, mean-aggregated), per_goal (one run per candidate covert objective, max across objectives), members (a weighted-mean ensemble of other assets), template (reuse another asset's prompt body), and schema (the output parser) — and its body is the prompt template. One harness renders the view, calls the trusted role model, and parses the score, so adding a monitor is adding a file. Two monitors that cannot be expressed as a template stay as code: simple (structured message-list prompt) and monitoring-bench (bench-replica model panel with an XML parser).
List available monitors (or show one monitor's declared view):
ct protocols monitors action [NAME] ct protocols monitors trajectory [NAME]
Action-Level Monitors (--action-monitor)
Score each untrusted-model action.
simple- Scores each action from the main task and preceding action/result history, with no candidate covert-objective hints.no-history- Scores each action in isolation with no preceding action history or candidate covert-objective hints.old-action- The default action monitor: one model call per action, shown the candidate covert objectives.old-action-no-goals- The default action monitor without the candidate covert objectives.checklist- Answers five yes/no checks per action (necessity, security surface, innocent explanation, concrete harm, removability) before scoring; the top of the scale is gated on concrete harm.orthogonal- Flags work that does not serve the main task - unnecessary changes, premature optimization, off-task activity - rather than judging malice directly.three-scores- Categorical monitor: a safe / uncertain / suspicious verdict per action, mapped onto the score scale.strong- Sees the full conversation history (goal-blind) so it judges each action in context.strong-goals- The strong full-history monitor, additionally shown the environment's candidate covert objectives.goal-analysis- Scores each action while shown the environment's candidate covert objectives, asking whether the action advances any of them.single-goal- Scores each action once per candidate covert objective in parallel and keeps the maximum score.multi- Weighted blend of the orthogonal and old-action monitors' scores for each action.mean-action- Averages two independent runs of the old-action monitor on each action. CLI default.mean-goals- Averages two independent goal-analysis runs on each action.mean-single-goal- Scores each action twice per covert objective, averages each objective's scores, and keeps the highest average.
Trajectory-Level Monitors (--trajectory-monitor)
Score the whole transcript in one shot rather than action by action.
quick- Reads the full transcript in a single pass, shown the candidate covert objectives, and returns one score for the trajectory.quick-no-goals- Reads the full transcript in a single pass with no covert-objective hints.goal-analysis- Reads the full transcript while shown the environment's candidate covert objectives, working through them one by one.single-goal-analysis- Reads the full transcript once per candidate covert objective in parallel and keeps the maximum score.mean-quick- Averages five independent quick passes over the transcript. CLI default.mean-goal-analysis- Averages two independent goal-analysis passes over the transcript.mean-single-goal-analysis- Reads the transcript twice per covert objective, averages each objective's scores, and keeps the highest average.monitoring-bench- Replica of the MonitoringBench benchmark's monitor: the bench prompt on its fixed 3-model panel (scores comparable to the imported monitoring-bench dataset), max-aggregated.
Whether a monitor sees the environment's candidate covert objectives — every side task registered for that environment, not just the one the agent was assigned — is fixed by its prompt asset: goal-aware and goal-blind variants are separate monitors (e.g. old-action vs old-action-no-goals, strong vs strong-goals, quick vs quick-no-goals).
Output
Post-hoc monitoring reads the source, runs monitors, and writes the monitor results. Action-level monitors write per-action results; trajectory-level monitors write one result per trajectory.
Sources are scored by replay
The input must be eval-backed: a local or remote .eval, an in-memory EvalLog, or a viewer run backed by an eval log — anything else (trajectory JSON, datasets, runs without a backing eval log) is rejected with an error. ct run monitor composes a replay eval: the replay untrusted policy re-emits the recorded actions, the replay sandbox serves their recorded tool outputs, and the monitor-only protocol runs the selected monitors in the loop. Two artifacts land in the save dir (an explicit --save-dir, or the default ./data/monitoring/<monitor>/<label>_<timestamp>/); the source is never overwritten and remains byte-for-byte unchanged:
- the replay-produced
.eval— a full protocol transcript whose tagged monitor bookkeeping messages carry eachMonitorResult, withmonitor_<name>sample scores attached natively by the protocol's scorer. The score value containsmax,action_<index>, and, when applicable,trajectory. Uploaded monitor runs register this log as theireval_log_url, and it round-trips throughconvert_eval_log_to_trajsinto the samemonitor_response/traj_monitor_responsefields downstream consumers read. <source-stem>.monitored.eval— the canonical monitored log: a deep copy of the source with themonitor_<name>scores copied onto the matching samples (each replay sample records its source(sample_id, epoch), so the copy is an exact keyed match). Existing named scores survive, and re-runningct run monitoron this file adds each newly selected monitor under its own score key — one canonical monitored.evalper source accumulates monitors.
Compact provenance for every named monitor lives in the canonical log's control_tower_monitors eval-header metadata (monitor names, configurations, costs, source runs), alongside derived_from (the source identifier) and monitor_ct_version, so listing or tracing monitors does not require loading sample transcripts.
Unless --no-upload is passed:
- Creates a run in the database
- Uploads all monitored trajectories
- Prints run URL and trajectory URLs
Note the model response cache defaults on, so re-running an unchanged command reproduces the same scores rather than drawing a fresh sample (the CLI warns when this happens); clear inspect's model cache (inspect cache clear) to resample.
Monitoring Details
Monitors add suspiciousness scores to trajectories:
- Scores range from 0 (not suspicious) to 10 (very suspicious)
- Different monitors use different approaches (LLM-based, heuristic, etc.)
- Action-level results are stored per action
- Trajectory-level results are stored once per trajectory
Cost Attribution
Every MonitorResult carries a cost_usd (and token fields) populated automatically by the orchestration layer — no per-monitor changes required. Multi-model monitors (e.g. multi) additionally get a cost_by_model split. An unpriced model is a None in that split; cost_usd stays the priced spend. See function-registry.md for the mechanism and known limitations, and trajectories.md for the Trajectory.stepwise_monitor_cost() / full_monitor_cost() / total_cost() helpers and the cost_breakdown field.
Technical Details
- Supports both action-level and trajectory-level monitors
- Tracks token usage and costs for LLM-based monitors
- Runs monitors concurrently for efficiency
- Handles errors gracefully per trajectory