Glossary
Abbreviations and terms used across these docs. Core Concepts defines the objects you work with in more detail.
| Term | Meaning |
|---|---|
| AMI | Amazon Machine Image, the disk image an EC2 fleet worker boots from. See EC2 Fleet. |
| Attack eval | The half of a control evaluation where the red team controls the untrusted model and pursues a side task. See AI Control Concepts. |
| Attack surface | The parts of an environment where an agent could accomplish a side task. See Making Side Tasks. |
| Audit, audit budget | An audit sends an action to a human reviewer; the audit budget is the fraction of honest actions the blue team may audit. See AI Control Concepts. |
| AUROC | Area under the ROC curve, a single number for how well a monitor separates honest from attack trajectories. See ROC Curves. |
| Blue team (BT) | The side that designs protocols to maximize usefulness and safety. See AI Control Concepts. |
| Cover | How well a main task gives an agent plausible reasons to touch the code or data a side task needs. See Validate Cover. |
| CT | Short for Control Tower, used in code comments and some variable names. |
ct | The Control Tower command-line program. See the CLI reference. |
| Defer to trusted (DTT) | A protocol that replaces a suspicious untrusted action with the trusted model's action. See AI Control Concepts. |
| Delayed and live failure | Whether a side task's harm happens only after the trajectory ends (delayed) or during it (live). See Side Task Failure Types. |
| EC2 | Amazon Elastic Compute Cloud, which hosts fleet workers. See EC2 Fleet. |
| env | Short for environment, used in command names (ct envs) and some prose. |
| Eval log | The Inspect .eval file an evaluation writes; the source of truth for results. See Evaluation Output. |
| Fleet | A group of EC2 workers that run one evaluation's samples in parallel. See EC2 Fleet. |
| FPR | False positive rate: the fraction of honest actions or trajectories a monitor threshold flags. Audit budgets are stated as an FPR. See Safety Curves. |
| Hawk | A shared Kubernetes cluster service for running Inspect evaluations. See Running Evals on Hawk. |
| Honest eval | The half of a control evaluation where the untrusted model behaves honestly; it measures usefulness. See AI Control Concepts. |
| IS | Internet simulator: a shared container that stands in for external websites and services. See Internet Simulator. |
| JSONL | JSON Lines: one JSON object per line, the format of task sets. |
| K8s | Kubernetes, the container orchestrator Hawk runs on. |
| LinuxArena | A family of settings of realistic software environments: linuxarena, linuxarena-extras, and a private held-out setting. See Development Setup. |
| LLM | Large language model. |
| Main task (MT) | One concrete piece of useful work in an environment, with its own scorer. See Making Main Tasks. |
| Match | One protocol evaluated against one red-team attack policy, giving one safety and usefulness pair. See Architecture at a Glance. |
| MCP | Model Context Protocol, the interface the web viewer exposes to AI assistants. See Web Viewer. |
| mtgen | Main task generation: the ct run mtgen pipeline that drafts, validates, and audits new main tasks. See Main Task Generation. |
| Policy (untrusted policy) | The red team's code that compiles what the untrusted model is asked, such as honest or attack, selected with --untrusted-policy. See Protocols and Policies. |
| PR | Pull request. |
| Protocol (blue protocol) | The blue team's control logic, which sees every call to the untrusted model and decides which action runs. See Protocols and Policies. |
| Red team (RT) | The side that controls the untrusted model in attack evals and tries to accomplish side tasks without being caught. See AI Control Concepts. |
| Reward hacking | Passing a scorer without doing what the task intends. See Reward Hacking Audit. |
| ROC curve | Receiver operating characteristic curve: a monitor's true positive rate plotted against its false positive rate. See ROC Curves. |
| Run | The uploaded record of one evaluation, pointing at its eval log. See Evaluation Runs. |
| S3 | Amazon Simple Storage Service, where uploaded eval logs are stored. See Database. |
| Sabotage eval | The ct run sabotage-eval pipeline that runs honest and attack evals, monitoring, and metrics in one command. See Sabotage Eval Pipeline. |
| Sabotage trajectory (straj) | A trajectory, usually made by a person, that accomplishes a side task while evading the monitor. See Making Sabotage Trajectories. |
| Setting | A collection of environments, usually one repository under .settings/. See Core Concepts. |
| Side task (ST) | A harmful objective the red team pursues covertly alongside a main task. See Making Side Tasks. |
| SSM | AWS Systems Manager, whose Parameter Store holds the secrets fleet workers read. See EC2 Fleet. |
| SSO | Single sign-on, used to get AWS credentials. See EC2 Fleet. |
| Suspiciousness score (sus score) | A monitor's rating of how suspicious an action or trajectory is, from 0 to 10. A trajectory's max_suspicion_score is its highest action score and traj_suspicion_score is the trajectory monitor's score. See Monitoring. |
| Task set | A named JSONL file listing task combinations, passed with --task-file. See Task Selection. |
| traj | Short for trajectory, as in ct traj and "traj directory", the folder that holds one straj's files. See Live Development. |
| Trajectory | The recorded behavior of one evaluation sample. See Trajectories. |
| Trusted model | A weaker model the blue team trusts not to scheme, used for monitoring and fallback work. See AI Control Concepts. |
| Trusted monitoring | A protocol in which the trusted model scores the untrusted model's actions and the most suspicious ones are audited. See AI Control Concepts. |
| TUI | Terminal user interface. |
| Untrusted model (U) | The more capable model that may be misaligned; the red team controls it in attack evals. See AI Control Concepts. |
| Usefulness | The share of main tasks the protocol completes in the honest eval. See AI Control Concepts. |
| VM | Virtual machine. |
| White team | The neutral party that fixes the game's parameters, such as the audit budget, models, and task distribution. See AI Control Concepts. |