Watch an expense agent through its first hour. It asks before opening the folder of receipts. It asks again before reading one of the PDFs inside. Then it asks, in the same flat sentence, before filing the claim to the accounting system and moving real money out of a real account. Three questions, one shape, one weight. You answered all three the same way, because nothing in them told you which one mattered. The agent was supervised the whole time. It was never actually overseen: not one of those approvals was calibrated to what the agent could break.

The mountain grades the run, not the skier

A ski resort does not run on a single rule. The runs are graded green, blue and black long before anyone arrives, and the grade describes the terrain — what it does to a person who gets it wrong — not the person standing at the top. On day one a new skier stays on the green with an instructor an arm’s length away. By the end of the week she takes the blue side alone and nobody follows her down, not because she has been declared trustworthy in general, but because a fall there is something you get up from. Meanwhile the black run under the cliff band is closed today, and it is closed for the instructor too.

Two scales, running at once. One moves with the person, one moves with the ground. Agent oversight works the same way when it works at all. A new user approves every single filesystem access. After a run of sessions that went well, on tasks that can be undone, that same user switches on auto-approval with a log and a button that stops the agent mid-action. And when the task is a high-impact change to infrastructure, the runtime pulls autonomy back down and demands an explicit escalation, no matter how many good sessions came before it.

The model is never the only defence

Five properties have to hold together for any of that to mean something, and they have to hold across the model, the harness around it, the tools it can call and the environment it runs in. Human control is the one you can see: which tools are switched on at all, what permissions they carry, whether a plan gets approved before it runs. Secure interaction is the one you feel only when it fails — model-side defences, monitoring, red teaming, and tools cut down to the least privilege that still does the job. Transparency is the boring one that decides whether anybody can reconstruct what happened: a trace of the decisions and of the limits the agent hit.

Underneath both sit the two nobody demos. Alignment is not a paragraph in the prompt telling the agent to be careful; it is training and evaluating the thing to recognise its own uncertainty and stop. Privacy decides what data the agent touches and how long any of it survives.

Put them on the expense agent and it looks like this. It shows the plan before sending anything. It asks about the ambiguous travel policy instead of guessing. Its tools hold reduced permissions, and every decision leaves a trace. So when one of those receipt emails turns out to carry a prompt injection, the model’s judgment is not the last line of defence. The tool boundary and the environment cap what a fooled agent can reach.

Autonomy is a reading, not a setting

You find out whether any of this is real by measuring it: how long the autonomous stretches run, how often people approve, how often they interrupt, how often the agent stops by itself to ask. Counting approvals alone proves nothing — a person who approves everything and a person who reads everything produce identical numbers. Neither can you read the risk of a whole session off one tool call stripped of its context. And telemetry that never states its own coverage, bias and privacy limits is a story, not a measurement.

The uncomfortable part is that every team is currently deriving these properties from scratch, one deployment at a time, which is exactly why shared benchmarks, shared evidence and open protocols matter more than any one company’s policy page. Autonomy is not a dial on the model. It emerges from the product, the user and the task, which is also where the human belongs in the loop. Oversight that treats every action alike measures nothing, and people route around it by the second week.