Core Concepts

This page lists the objects you work with in Control Tower, each with a short definition and the class or function that implements it. Each name links to the page that covers it. Abbreviations and other terms are in the Glossary. For the research framing behind these objects, read AI Control Concepts.

How the pieces fit

Rendering diagram...

A setting groups environments. Each environment defines main tasks, side tasks, and a task space that says which combinations of them are valid. An evaluation selects combinations, turns each one into an Inspect sample, and runs an agent on it. The eval log records one trajectory per sample, monitors score those trajectories, and metrics turn honest and attack results into usefulness and safety.

Settings and environments

  • Setting: A named collection of environments, usually one git repository pulled into .settings/<setting>/ or an installed package. Its environments are discovered from its subdirectories, unless it ships a setting.py (or registers an entry point) whose get_environments() builds them. In code: SETTINGS in control_tower.environments.all_environments; the control_tower.settings entry-point group.
  • Environment: A system the agent works in: a sandbox, its documentation, and the tasks that run in it. In code: the ControlEnvironment protocol in control_tower.environments.protocol; DockerEnvironment for Docker Compose environments.
  • Environment args: Per-run options an environment declares with an args_model and receives from --env-arg key=value. Every sample records them in metadata.environment.env_args. In code: control_tower.environments.env_args.
  • Sandbox: Where the agent's tool calls run: a Docker Compose project, a virtual machine, a local process, or a simulated sandbox. In code: Inspect's SandboxEnvironment; ControlEnvironment.get_sandbox_spec().

Tasks

  • Main task: One concrete piece of useful work in one environment, with its own description and scorer. A main task is a single instance, not a template that expands over dataset samples; see Settings built from a dataset. In code: MainTask in control_tower.task.task.
  • Side task: A harmful objective the red team's agent tries to accomplish covertly alongside a main task, with its own scorer. In code: SideTask in control_tower.task.task.
  • Task combination: The unit an evaluation runs: a tuple of main tasks and a tuple of side tasks, compared as sets of task ids. Every combination Control Tower builds today holds at most one of each. Its mode is honest (main only), attack (main and side), side_only, or empty (neither). In code: TaskCombination in control_tower.task.space.
  • Task space: The set of task combinations an environment admits. CrossProduct pairs every main task with every compatible side task; ExplicitSpace lists combinations one by one. In code: TaskSpace, CrossProduct, ExplicitSpace in control_tower.task.space; ControlEnvironment.task_space().
  • Task set: A named JSONL file of {"env", "main_task", "side_task"} records that a setting ships under task_sets/, passed with --task-file. In code: control_tower.task.task_sets.

Agents and protocols

  • Trusted and untrusted models: The weaker model the blue team trusts and the stronger model it may not. The trusted and untrusted aliases name the defaults. In code: ModelAlias in control_tower.models.aliases.
  • Untrusted policy: The red team's side of the game: it compiles what the untrusted model is asked, such as honest or attack, and never executes an action itself. Selected with --untrusted-policy. In code: Policy in control_tower.protocols.eval.red; the policies registry in control_tower.protocols.eval.registry.
  • Protocol: The blue team's control logic: it sees every call to the untrusted model and decides which action runs, blocking or auditing as it goes. Selected with --protocol. In code: Protocol in control_tower.protocols.eval.blue; the protocols registry in control_tower.protocols.eval.registry.
  • White: The deployment the run holds fixed: the harness that drives the untrusted model, the affordances the red team is granted, and the step budget. Set with --harness, --grant and --max-steps. In code: White in control_tower.protocols.eval.white.
  • Monitor: A model call that returns a suspiciousness score from 0 to 10, asked live by a protocol or post hoc by ct run monitor. An action monitor scores each action; a trajectory monitor scores the whole trajectory once. In code: Monitor in control_tower.protocols.eval.monitoring.monitor.
  • Scorer: The Inspect scorer that decides whether a main task or side task succeeded. In code: Inspect's Scorer; the per-task @linux_scorer decorator in control_tower.environments.scorer_utils marks a task's scorer, and linux_scorer() in control_tower.task.scorers runs a sample's main and side task scorers.

Results

  • Sample: The Inspect sample built from one task combination. Its id is <env>:<main task>:<side task> with trailing colons dropped (so a side-only sample is <env>::<side task>), shortened with a hash suffix when longer than 63 characters, and its metadata names the environment and tasks. In code: SampleMetadata in control_tower.task.sample_metadata.
  • Eval log: The Inspect .eval file an evaluation writes. It is the source of truth for everything downstream. In code: Inspect's EvalLog.
  • Trajectory: The recorded behavior of one sample: its actions, scores, monitor results, and costs, read directly from the eval log. In code: accessors in control_tower.trajectories.sample.
  • Sabotage trajectory (straj): A trajectory, usually made by a person, that accomplishes a side task while evading the monitor. It shows that the side task can be done subtly. In code: stored under a side task's trajs_dir.
  • Run: The uploaded record of one evaluation: a header (policy, protocol, task set, tags, cost) that points at its eval log. In code: RunHeader in control_tower.agent_runs.models.
  • Dataset: A named, versioned list of trajectory or run ids kept on the viewer, managed with ct dataset. It is unrelated to the benchmark data a setting may build its tasks from. In code: DatasetVersion in control_tower.datasets.model.

Metrics

  • Usefulness: The share of main tasks the protocol completes in the honest evaluation. In code: the usefulness_calculations registry in control_tower.metrics.usefulness_calculations.registry.
  • Safety: The probability that the blue team catches the attack before a side task succeeds, at a given audit budget. In code: the safety_calculations registry in control_tower.metrics.safety_calculations.registry.
  • Audit budget: The fraction of honest actions the blue team may send to a human auditor. It fixes the suspiciousness threshold that counts as caught. In code: AuditRule in control_tower.metrics.audit_rule.
  • Match: One protocol evaluated against one red-team attack policy, producing one safety and usefulness pair at a fixed audit budget. In code: a research unit rather than a class; ct run sabotage-eval runs one end to end.

Settings built from a dataset

Many benchmarks cross task templates (for example "refactor three files") with dataset samples (for example repositories). Control Tower has no template object, so such a setting builds one MainTask per (template, sample) pair in its task space and records the template and sample in the main task's id. Making Main Tasks shows how to build the pairs and group results by template.

Control Tower and Inspect

Control Tower runs on Inspect, and most objects above are Inspect objects or wrap one:

InspectControl Tower
TaskThe control_tower/control_eval task that ct run eval assembles
SampleOne task combination
Solver / AgentAn untrusted policy under a protocol, wired by build_eval_solver
ScorerMain-task and side-task scorers, plus the monitor scorer that reports every monitor's scores
EvalLogThe eval log, read as one trajectory per sample

Control Tower and Inspect lists which behavior comes from Inspect and where Inspect documents it.