Most AI pilots die after the demo, not during it. The missing piece is usually treated as a disclaimer. The approval gate is the product that makes the agent shippable.
In the room, the agent drafts the reply, updates the spreadsheet, or files the ticket. Someone nods. A week later nobody can say who owns the workflow, which accounts the agent may touch, or what happens when the draft is wrong. The model did not fail. The system never existed.
The missing piece is usually treated as a disclaimer: “of course a human will review it.” That sentence is doing the real work. The approval gate is not a footnote on an impressive agent. It is the product that makes the agent shippable.
Why demos skip the gate
A live demo rewards fluency and speed. A pause for “does this look right?” feels like friction. So the builder removes it, or puts it on a slide and never builds it. Then production arrives with real customers, real money, and a real CRM. Suddenly every write needs a person, and the “autonomous” agent becomes a queue of drafts nobody has time to triage.
That is not a model problem. It is a design problem.
What a useful gate actually does
A good review step is not “a human stares at everything.” It is a narrow checkpoint on consequential actions, with enough context to decide quickly.
Write down four things before you pick a model:
- What counts as a write. Sending email, charging a card, editing production records, installing software, changing permissions. Reads and drafts can be freer. Mutations need a yes.
- Who may approve. One role, not “whoever is free.” If that person is out, name the backup. An agent without an owner is a liability with a chat window.
- What the reviewer sees. The proposed action, the source of the data, and the one-line reason the agent thinks this is correct. Without that, review is theater.
- What happens on “no” and on silence. Refuse, queue, or escalate. “Wait forever” is how tickets rot.
If you cannot answer those four, you are not ready to put the agent on a real account. You are ready for a sandbox.
Keep the gate from becoming a bottleneck
Teams fear review because they imagine approving every token. That is the wrong shape.
Batch routine drafts into a short queue with a clear SLA (for example, end of day). Auto-approve only the cases you have scored on an eval set and that stay inside a hard boundary, same template, same destination list, no payment fields. Escalate the edge cases. Measure how long a review takes; if it regularly exceeds a few minutes, the agent is handing over the wrong artifact.
The goal is not zero human time. The goal is human time spent on decisions that matter, not on reconstructing what the agent already knew.
Same discipline for automations and custom software
This is not only an “agent” story. A productized automation that drops finished copy into a folder still needs a publish step if the brand cannot afford a bad send. An internal tool that replaces a spreadsheet still needs roles and permissions so the intern cannot overwrite last quarter’s numbers.
At Viking Labs we treat the boundary as part of version one: what the system may do alone, what requires approval, and what waits for version two. ResearcherFlow is the production case we run ourselves, live billing, permissions, APIs, and a research agent that can draft notes and tasks but does not write into the workspace until a human approves. That is the same discipline we bring to a client workflow that is currently stuck in Slack and a shared drive.
A one-week test
Pick one painful, repetitive job. Write the definition of done and the four gate answers above. Build the smallest path that produces a reviewable draft, not a fully autonomous loop. Run ten real cases. Score them. Only then widen the scope.
If the gate feels annoying, fix the handoff, better context, clearer options, fewer false positives, before you remove the gate. Removing it to “feel autonomous” is how you get another demo that cannot survive contact with production.
Bottom line
The impressive part of an AI system is rarely the model release notes. It is the boring control plane: an owner, scoped access, an evaluation set, and a human yes before anything consequential leaves the building.
That is what we build. If you have one workflow that is ready for that treatment, map the first build: tell us the workflow
Jon Marrs · About Viking Labs
Leave a Reply