Protocols and Policies

An eval is a game between two teams under a deployment a third one holds fixed (see AI Control Concepts). The code for all three lives in control_tower.protocols.eval:

  • Red is the untrusted policy (--untrusted-policy): a Policy that compiles what the untrusted model (U) is asked. Its registry is policies, listed by ct protocols policies.
  • Blue is the protocol (--protocol): a Protocol that sees every call to U and decides which action runs. Its registry is protocols, listed by ct protocols protocols.
  • White is the deployment (White, set by --harness, --grant and --max-steps): which harness drives U, which environment affordances red is granted, and how many actions U may take.

CLI Usage

ct run eval --untrusted-policy honest --task-file simple ct run eval --untrusted-policy attack --model-role untrusted=sonnet --task-file medium ct run eval \ --untrusted-policy simple-attack \ --protocol trusted-monitoring \ --protocol-arg 'monitors=[simple]' \ --protocol-arg sus_threshold=7 \ --task-file simple

--untrusted-policy defaults to honest and --protocol to null-blue-team, which runs every action unmodified. ct protocols policies <name> and ct protocols protocols <name> show an entry's arguments. Policies and protocols defined in your own project are selected by module:function (e.g. --untrusted-policy my_project.policies:my_policy); see Function Registry.

Package Layout

  • red.py: Policy, and Untrusted, the controller that owns U's action model and granted affordances and drives red's private calls.
  • blue.py: Protocol, Audited, and the Game that routes every call to U through the protocol.
  • white.py: White and the harnesses (NullHarness, InspectSweHarness).
  • game_view.py: BlueView and RedView, what each team is told about a sample. Blue's view carries no side task, not even whether one exists.
  • solver.py: build_eval_solver, which wires a policy, a protocol and a White into an Inspect solver.
  • policies/, protocols/: the registered policies and protocols.
  • monitoring/: the Monitor model, its packaged prompts, and the monitor scorer.

Policy Design

A policy is a function registered with @policies.register() whose parameters are its arguments, returning a Policy. A Policy holds no model and no affordances. It implements two methods:

  • frame(input) returns the conversation U is asked. input already carries red's private turns, tagged ct_private in metadata, for the policy to keep, move, summarize or drop.
  • offered_affordances(granted) states which of white's granted affordances U may call as private tools; super() offers all of them.

setup(view) builds per-sample state from the RedView, including the side task. probe(name, **kwargs) runs one granted affordance through the controller, unseen by blue. A policy that is not a model (such as replay, test or sleep) overrides action_model().

Protocol Design

A protocol is a function registered with @protocols.register() returning a Protocol. It makes two decisions:

  • on_input(view, conversation) may rewrite what U is about to be asked, or return None to leave it; super() leaves it.
  • on_output(view, draw) asks draw(n) for n candidate actions and returns the one to run, or Audited(draft) to send that draft to a human and end the sample.

self.block(call, reason) refuses one tool call: the deployment runs the engine's blocked tool in its place and reports reason. Protocols read the actions that already ran with untrusted_conversation(transcript().events). Monitors are Monitor values passed as a protocol argument (a packaged prompt name, a mapping of Monitor fields, or {monitor: module:function}); a protocol scores a draft with score_all, and the scores land in the sample store, where the monitor scorer reports them. See protocols/trusted_monitoring.py for a complete example.

Metadata

An eval's log records untrusted_policy_selection and protocol_selection (name and arguments as given), the white it ran under, the model each role was bound to, and eval_version. Trajectory conversion stores the team choices in trajectory metadata as untrusted_policy and protocol.