Approve the Decisions. Watch the Work.
What to put your hand on while an agent runs, what to keep your hand off, and how the same line decides when it runs without you.
Everyone has worked for the manager who needs a status update every hour. Nothing ships without him seeing it first. He's not malicious. He just can't let go of the work itself. And the whole team slows to the speed of his attention, because that's the actual bottleneck now. Him.
That manager is how most people supervise an AI agent. They wire an approval into every step, sign off again and again, and then wonder why the loop they built to save a day still eats one. They rebuilt the bottleneck. They just moved it from doing the work to approving it.
The question I keep getting is when do you stop checking the agent. Wrong question. The real one is what you put your hand on while it runs.
There are two ways to get this wrong. Approve every step, and you’re the manager above, the bottleneck. Approve nothing, and one run does something you can’t take back. The fix isn’t the midpoint, gating every call the agent makes. Half of them it should just make on its own. You keep your hand on two: what you can’t undo, and the call that’s the actual point. Everything else, the work and the small reversible choices along the way, runs without you. You watch it. You don’t gate it.
The calls that are yours
Two of the agent’s calls are yours. The one that’s the actual point: picking what’s worth doing. And the one you can’t take back. Those you gate, every time. The rest, the reversible calls it can make and you could undo, are the agent’s. You let them go.
Here's one I run with AgentUse. It turns my queued quote posts into images and ships them to X, Facebook, and Instagram, and it stops in exactly one place: before anything publishes. The agent checks what's ready, builds the post, finds a slot, all on its own. Then it holds, because pushing something to a public account is the step it can't take back. That gate is mine, and it sits exactly where the irreversible thing happens.

A loop can have more than one of these, and they don’t all sit at the end. My newsletter stops twice: once in the middle, the moment it’s laid out the week’s candidate angles, for me to pick the one worth saying, and again at the end for me to approve the send. The first one is mid-run, the agent’s already done a chunk of work before it stops, and that’s fine. A mid-run gate isn’t the problem. The gate belongs wherever one of these calls lives, and a call can land anywhere in the run. One agent split by three approvals is clean if each one is a call that’s actually yours.
The test is never where the gate falls in time. It’s this: is this a call only I should make, the irreversible or the essential, or am I babysitting work the agent should just do?
The execution isn't
The execution is everything between the decisions. The drafting, the formatting, the assembly, the calls to tools. None of it needs your judgment, and every approval you staple onto it is just you in the way.
This is the manager from the top of the piece. An employee who stops every twenty minutes to ask "is this okay so far?" isn't being managed. He's being blocked. And he never builds the judgment to run without you, because you never let a single stretch of work finish without your hand on it. Gate the execution and you cap the loop's throughput at your availability, which was the one thing you were trying to escape. You also tell the agent you don't trust it to execute, only to start and stop, which is another way of saying you don't trust it at all.
So when you feel the urge to add a gate, ask what it’s sitting on. One of your two calls, the irreversible or the essential, keep it. Anything else, a step or a choice the agent can undo, and you’re rebuilding the bottleneck.
Watch the work, don't block it
Keeping your hand off the execution does not mean going blind. This is the other ditch people drive into. They think the only two options are approve every step or ignore everything until the end.
There's a third, and it's the whole game: watch without stopping. Observability. I replay exactly what the agent did and why, step by step, after the run, without ever blocking the run. I'm reading its real reasoning, not its final summary. Here's the read-only agent that closes the loop on my SEO pipeline, a run I never watched live: it pulled the search funnel, hit a command it wasn't allowed to run, worked around it, and wrote its snapshot. Every step, and the reason behind each one, is right there whenever I want it.

That's what a good manager actually does. He doesn't approve every paragraph his writer drafts. He reads the work, keeps a feel for how it's going, and steps in at the decisions. Observability is that, for agents. It's how you supervise the work without standing in it.
And it changes what the review even is.
The review is a decision, on evidence
Approving the output is itself a decision, the last one, so it gets a gate. But the gate isn't a thumbs-up on a final answer floating free of context. It's the output plus the trace of how the agent got there. I approve the rendered newsletter, the real email as it will land, and I can see the run behind it. Evidence, not vibes.
Whether that gate ever comes off depends on one thing, and it isn't how long the agent has behaved. It's what a mistake costs and whether you can take it back.
Reversible and cheap: once the replay has shown me enough clean runs, I drop the gate and let it ship on its own, and I keep watching the trace. Irreversible and expensive: I keep my hand on it for as long as the company exists. The newsletter send goes to the whole list and can't be unsent, so I press send myself, forever, and it has nothing to do with the agent being unreliable. It's the most reliable thing I run. The support agent drafts replies and never sends the refunds and cancellations, because a wrong one is a real person and real money I can't claw back. Drafting is reversible. Sending is not. That line, not the agent's track record, decides which decisions stay gated.
The arc is just onboarding
Put it together and it's the shape of training anyone.
Week one, you brief the new hire, you read everything they produced, you approve before it goes out. A month in, you brief them, you skim, you only sign off on the things that can't be undone. A year in, you brief them and the only thing that reaches you is the exception they flagged themselves, because they learned to raise a hand on the weird cases instead of guessing.
The agent is the same arc. What changes isn't the agent. It's you moving from reading every trace to trusting that the trace exists and watching the number. An agent that never escalates isn't reliable, by the way. It's just quiet about being wrong. The hand-raise is the thing you're really training.
And it runs backward too. The week the replay starts showing sloppy steps, or one bad output slips the last gate, I go back to reading every run and the gate moves back up. Trust is revocable because the evidence is revocable. Nothing about last month protects you when the model under it shifted or the inputs drifted and the instructions didn't.
So stop asking how much you trust the agent. Ask what you’re holding. The two calls that are yours, on the evidence of how it got there. And nothing else. Watch the work the whole time, and never once stand in it.
The judgment was always the job. The work was never yours to hold.




