Businesses that gave agents the long, messy jobs are hitting a ceiling. The software outruns any manager who is supposed to check it, and it does not wait. That is an oversight failure, not a flashy demo going sideways.

Travel booking is the cute version. Production systems, customer records, or money is the real one. TechCrunch laid out the bind as speed and duration at a volume people cannot realistically review.

Human review dies at agent volume

Human-in-the-loop was the safety story of the first chatbot wave. A model drafted. A person sent. That story cracked when the software learned to use tools: browsers, calendars, terminals, payment APIs. Once an agent can take actions, the bottleneck is no longer generation. It is attention.

Security teams already know this shape. You do not have enough analysts to read every log. You write detectors, then you write detectors for the detectors, and you still miss things. The difference is that those tools were not also the intern booking flights. Agent products are both the worker and the thing that needs watching.

By 2025, coding agents and computer-use demos had made the volume problem obvious. A person can review a pull request. A person cannot review four hundred of them that arrived while they were in a meeting, especially if the agent is also opening tickets and paging the on-call. Vendors sold leverage. Reality invoiced supervision.

Companies did not stop deploying. They asked for a scaler. The scaler, in the current conversation, is more AI. Systems that watch other systems. It has an obvious circular quality. It may also be the only scale that matches the workload.

Watchers watching watchers

A second model that flags odd tool calls, odd spends, or odd data access is not a person. It is a filter. Filters are how every high-volume system copes, from credit cards to content moderation. They also fail in batches.

Human review does not disappear in this model. It gets reserved for the cases the watcher marks, plus whatever audit sample a compliance team still demands. That is how anti-money-laundering software already works. It is also how you sleepwalk into a regime where the only things a person sees are the things another model thought were interesting.

Whether that actually catches bad behavior, or just adds another layer that can fail, is not settled. A watcher trained on the same stack as the worker can share the worker's blind spots. A watcher from a different vendor adds cost and latency, plus a new contract. Neither option is a philosophy. Both are procurement.

If the watcher is also an agent, it can be talked into looking away. That is the same jailbreak problem the industry has had since chatbots, now sitting on top of a system that can act. Oversight you can negotiate with is not oversight. It is a conversation.

Consumer products are not immune to the same more software on top of software reflex. Roku stacking 30-plus streaming bundles is a mild version: another layer meant to simplify a mess it also helped create. Agent oversight is the sharp version. Get it wrong and you did not miss a show. You shipped a refund, a bad deploy, or a leaked inbox.

Phone buyers already know what it feels like when a vendor sells a new layer as the fix for last year's layer. The iPhone 18 Pro versus Pixel 11 Pro price fight is a $100 argument about hardware. This is an argument about whether you still have a human in the path at all.

What to demand before an agent touches production

Do not let an agent issue refunds, change IAM, or merge to main without a watcher log you can actually read. Demand a second-system trail, not a dashboard adjective. Prefer a watcher from a vendor that is not also selling you the worker, even if that costs more. If they are the same company, you are buying a mirror and calling it a guardrail.

Ask for a number: how fast a human has to see an exception, and who is on the hook when the watcher was wrong. Audit teams and insurers will ask for that number before your CEO does. Copy their checklist.

Short demos never showed the review problem. Long jobs do. The firms giving agents multi-step research, repo-wide refactors, travel that spans three systems, or support that can issue money are hitting this first. Everyone else meets it the week their vendor ships autonomous as a checkbox.

I would not run unattended agents on anything you cannot undo. I would run a watcher on anything you do run, and I would still sample the watcher's misses. The iPhone 18 Pro Max's familiar shell is a reminder that a new layer on an old body is not automatically a new product. Agent oversight is the same trick. Make them prove the new layer catches something a person would have caught, on a clock a person could not have kept.