Watching the Watcher
Your monitor has a clean record. It reads every trajectory the agent produces, it flags the tool calls that deserve flagging, and in six months of production it has not missed anything you know about. Then an outside auditor runs a session against it: a dozen ordinary-looking tool calls, each one defensible on its own, that together move a customer list somewhere it was never meant to go. Nothing fires. The run finishes clean, the way every run has finished clean. The record was never evidence that the monitor worked. It was evidence that nobody had tried it.
A door is rated by its attackers
A strongroom door is not rated by the people who welded it. It goes to a test house, where a team is paid to get through it, with a clock running that nobody stops for lunch. They bring drills, angle grinders, a cutting torch, and what they write down is not an opinion: these tools, in this order, this many minutes to open a hole big enough to reach through. That figure is the rating. And when somebody finds a faster way in — a weak seam near the hinges, a plate that peels once you have it started — the technique does not stay in that room. It goes into the protocol every later door is tested against, so that nobody has to remember it.
An adversarial red team is that crew, pointed at the part of your system whose job is to watch. Independent evaluators get controlled access to the boundaries that matter, and then they build trajectories on purpose: hidden attacks, evasion strategies, the edge cases your own tests were never shaped to look for. Each vulnerability is reproduced, classified by impact, and handed back with the evidence attached. Then the traces become a regression suite. That last step is the whole difference between a test and a show, because an exercise whose findings evaporate when the report is filed has proved nothing, twice.
A rating covers only what was tried
What comes back from the test house never says the door is secure. It says which tools were allowed, how long they took, and what was on the bench — the door, usually; the frame, sometimes; the fixings into the brickwork, often not at all. An insurer reads those lines before deciding how much cash you may keep behind it, and two doors from two makers can be held side by side because the fields are the same fields.
A safety transparency record is that paper, written for a deployed agent. For each release you set down who owns it, how much autonomy it has, which tools it can reach, which data and credentials sit inside its boundary, what safety measures exist, what testing outsiders did, and which incidents have already happened. Then you mark each line for what it is: declared by you, verified by somebody who is not you, or unknown. Before a procurement agent goes near a supplier, that list is what somebody signs. The failure here is quieter than lying. It is publishing capability claims with the limits left off, or letting your own testing sit under a heading that reads like independent validation. An empty field means nobody looked. It does not mean there is nothing there.
The fields nobody thinks to fill in
Some of what belongs in that record is not about the agent at all. A browser tool makes the point, because it is the kind of thing you install without reading. The server driving the browser can send usage statistics of its own — how often calls succeed, how slow they are, what environment they ran in — and it can do that by default, which means keeping it is a decision you never made. Switching it off is an explicit line of configuration, the sort of flag a team sets in CI so that test-environment data stays home. The mistake worth naming is assuming that disabling the browser’s own telemetry disables the tool server’s. Those are two flows, and only one of them is in the picture you are looking at.
None of it survives neglect, either. The record describes one model, one set of tools, one deployment boundary, one policy, and the moment any of those changes it is a claim about a system that no longer exists. Safety work famously does not compress into a score the way a capability benchmark does, and it never will. Comparable fields, honestly marked, are what you get instead, and they are worth more than a number, because somebody who is not you can check them.
A monitor nobody has attacked has a clean record, not a proven one.