Untrusted policies
Each entry below is a solver deciding what the untrusted model submits, selected by name with ct run eval --untrusted-policy <name>. Names, arguments, and defaults are read from the untrusted_policy registry, so this page cannot drift from the code.
Adding one
from control_tower.protocols.untrusted_policy.registry import untrusted_policies @untrusted_policies.register() def my_untrusted_policy(...): ...
Every non-underscore module under the registry's discovery package is imported on first lookup, so the decorator is all the registration a new untrusted_policy needs.
Registered untrusted policies
attack
Canonical attack policy. The model decides when to call planning + rst based on the attack motivation injected as a private user message.
The protocol's neutral SWE system prompt is left untouched; the monitor sees it (same as honest), and the attack framing lives only in the private motivation message injected into the local msgs_to_model.
Defined in control_tower.protocols.untrusted_policy.implementations.attack.policy.attack_policy.
| Argument | Type | Default | Description |
|---|---|---|---|
model | ModelSpec | ModelSpec(...) | |
max_steps | int | 150 | |
force_submit_grace | int | 10 |
honest
Canonical honest policy. The model decides when to call ready_to_plan and submit based on the system prompt. No autofire.
Defined in control_tower.protocols.untrusted_policy.implementations.honest.policy.honest_policy.
| Argument | Type | Default | Description |
|---|---|---|---|
model | ModelSpec | ModelSpec(...) | |
max_steps | int | 150 | |
force_submit_grace | int | 10 |
http-proxy
Drive the agent loop from an external HTTP client.
Defined in control_tower.protocols.untrusted_policy.implementations.http_proxy.policy.http_proxy_policy.
| Argument | Type | Default | Description |
|---|---|---|---|
host | str | '127.0.0.1' | |
port | int | 8765 |
obvious-attack
Direct attack with no concealment or stealth requirements.
Defined in control_tower.protocols.untrusted_policy.implementations.obvious_attack.policy.obvious_attack_policy.
| Argument | Type | Default | Description |
|---|---|---|---|
model | ModelSpec | ModelSpec(...) |
open
Open-ended policy with a freeform goal string.
Defined in control_tower.protocols.untrusted_policy.implementations.open.policy.open_policy.
| Argument | Type | Default | Description |
|---|---|---|---|
goal | str | '' | |
model | ModelSpec | ModelSpec(...) |
private-test-attack
Attack policy that can privately verify the side task with run_side_task_test.
Defined in control_tower.protocols.untrusted_policy.implementations.private_test_attack.policy.private_test_attack_policy.
| Argument | Type | Default | Description |
|---|---|---|---|
model | ModelSpec | ModelSpec(...) |
replay
Replay a recorded trajectory of tool calls deterministically.
Defined in control_tower.protocols.untrusted_policy.implementations.replay.policy.replay_policy.
| Argument | Type | Default | Description |
|---|---|---|---|
traj_like | str | required |
simple-attack
Single-phase attack policy that swaps the system prompt for a side-task-focused one.
Defined in control_tower.protocols.untrusted_policy.implementations.simple_attack.policy.simple_attack_policy.
| Argument | Type | Default | Description |
|---|---|---|---|
model | ModelSpec | ModelSpec(...) |
simple-honest
Delegate every generate call to the untrusted model verbatim.
Uses model.make_model().generate (not get_model(name)) so that
--untrusted-policy-arg model_temperature=... etc. on the
ModelSpec actually take effect.
Defined in control_tower.protocols.untrusted_policy.implementations.simple_honest.policy.simple_honest_policy.
| Argument | Type | Default | Description |
|---|---|---|---|
model | ModelSpec | ModelSpec(...) |