Web Development • HIPAA-Aware Healthcare Web • SEO • AI Search Optimization (407) 409-8383   |   [email protected]
Human in the loop · AI oversight · AI adoption

Companies are moving work to AI. The human in the loop is the point.

AI adoption is near-universal and mostly augments people rather than replacing them. What companies actually hand to AI, the evidence for keeping a human in the loop, and what that person really does.

By NavoTech  ·  Updated  ·  10 min read

Depending on which number you trust, almost every company is using AI, or barely one in five is. Stanford's 2025 AI Index found that 78 percent of organizations reported using AI in 2024, up from 55 percent the year before. The U.S. Census Bureau, which surveys a nationally representative sample of actual firms rather than a self-selected panel, puts business AI use closer to 18 percent, and concentrated in large employers: around 37 percent of firms with 250 or more employees, against under a fifth of the smallest ones. Both are true. The headline figure counts anyone who has touched AI in any function; the government figure counts firms that have genuinely put it to work. Either way the direction is the same, and it is up.

The part that gets lost in the adoption race is what happens to the people. The same Census data is blunt about it: AI is augmenting workers far more than replacing them. Augmentation is the primary effect at 44 percent of AI-using firms, and among firms that report any task-level effect, two thirds augment exclusively, while pure substitution shows up in only about 5 percent. So the realistic picture is not a person handed a pink slip. It is a task handed to a model, and a person whose job has quietly shifted from doing the task to running the thing that does it. That shift is the subject of this piece, and it is the part most vendors skip. Our own view, which shapes how we build AI systems and automation, is that the human in the loop is not a courtesy. It is the load-bearing part.

What companies are actually handing to AI

The tasks moving to AI share a shape, and the data shows it. At the worker level, the Census found generative AI is used most for document-style work: writing and editing at 85 percent of firms, with information search a distant second at 50 percent. By business function it clusters in sales and marketing, used by 52 percent of AI-using firms, then strategy, IT, and R&D. The accuracy-sensitive, judgment-heavy tasks, like customer support and software debugging, lag noticeably, and the government authors say plainly that this is likely due in part to "a continued requirement for rigorous human oversight." That is the tell. The work that has moved fastest is the work where a wrong draft is cheap to catch and fix.

That matches what makes a task worth automating. The builds that pay off are high volume, repetitive, and tolerant of a review step, which is exactly the profile we lay out in when AI automation actually pays off. First-draft generation, document and data extraction, intake triage, classification, and summarization all fit. It is why the useful projects tend to be workflow and document automation and assistants grounded in your own content, not a system asked to make final calls alone. The volume is what makes the build worth paying for. The review step is what keeps the occasional bad output away from a customer or your books.

On the right tasks, a person plus AI beats either alone. On the wrong ones, AI makes them worse.

The best evidence here is a field experiment run with the Boston Consulting Group and researchers from Harvard, Wharton, and MIT, Navigating the Jagged Technological Frontier. Across 758 consultants and 18 realistic tasks that sat inside what AI does well, the consultants using AI completed 12.2 percent more tasks, finished them 25.1 percent faster, and produced work rated more than 40 percent higher in quality. The lift was largest for the weaker performers, who improved 43 percent against their own baseline while the strongest improved 17 percent. On those tasks this is not a subtle effect. It is the reason companies are moving.

Then the same study did what most vendor demos never will. It handed people a task deliberately chosen to sit just outside the frontier, where the AI looks equally confident but is actually wrong. There, the consultants using AI were 19 percentage points less likely to reach the correct answer than the ones working without it. The model gave no signal that it had crossed a line; it produced a fluent, plausible, wrong answer, and the people who trusted it followed it off the cliff. The authors named the problem well: the frontier of AI capability is "jagged," an uneven edge you cannot see from the inside. Knowing which side of that edge a given task sits on is not something the model can tell you. It is the human's job, and it is the single most valuable thing the person in the loop does.

You see the augment-not-replace pattern in live operations too. In a study of 5,179 customer-support agents, economists at the National Bureau of Economic Research found that access to a generative-AI assistant raised issues resolved per hour by 14 percent on average, and by 34 percent for the newest and least experienced agents. The tool, in the authors' own words, "is designed to augment agents, who remain responsible for the conversation and are free to ignore its suggestions." Responsible, and free to ignore. That is the loop working as intended.

Even grounded on your own content, it is not reliably right

The reason a person cannot simply sign off unread is that these systems do not reach zero errors, even when you do everything right. Vectara maintains a public hallucination leaderboard that measures how often a model introduces unsupported facts when summarizing a document it was handed. The best model on it sits near 1.8 percent, which sounds excellent until you look at the harder, long-document version, where several of the strongest current models exceed a 10 percent hallucination rate. And that is the easy case, where the source is right there in front of the model. On real professional work the numbers get worse: a Stanford study of commercial, retrieval-based legal-research tools found they still hallucinated between 17 and 33 percent of the time, and the authors concluded that vendor claims of eliminating hallucinations are overstated.

This is why grounding a model on your material, the technique behind a good assistant on your own content, cuts errors sharply but never promises perfection, and why we build citations and a human handoff into those systems from the start. A review step is not evidence that the automation failed. It is the design that makes a good-fit automation safe to ship. Treat the output as a fast, capable first draft that a person confirms, especially for anything customer-facing or money-moving, and the error rate stops being a hidden liability and becomes a managed cost.

When the human is missing, the company still owns the answer

If you needed a single reason to keep a person accountable for what the AI says, a Canadian tribunal supplied it. In Moffatt v. Air Canada, decided in February 2024, the airline's website chatbot told a grieving customer he could claim a bereavement discount retroactively after booking. That was wrong, it contradicted the airline's actual policy, and Air Canada argued it should not be liable because the chatbot was, in effect, a separate legal entity responsible for its own actions. The tribunal called that "a remarkable submission," held that the airline is "responsible for all the information on its website" whether it comes from a static page or a chatbot, and ordered it to pay. The bill was small, about 812 Canadian dollars. The precedent was not.

The courts have been just as clear when professionals let the tool do their thinking. In the widely reported Mata v. Avianca case, a federal judge sanctioned two lawyers 5,000 dollars for filing a brief full of citations that ChatGPT had invented, writing that there is "nothing inherently improper" about using a reliable AI tool but that attorneys keep a "gatekeeping role" and had "abandoned their responsibilities" by not checking. The through line is simple and it does not bend: the AI's output is your output. Accountability does not transfer to the model, and no tribunal is going to let it.

Oversight is becoming a requirement, not a preference

What used to be good practice is turning into a written standard. The European Union's AI Act, in Article 14, requires that high-risk AI systems be "designed and developed in such a way ... that they can be effectively overseen by natural persons" while in use. The oversight it describes is specific: the assigned person must be able to understand the output, to "remain aware of the possible tendency of automatically relying or over-relying" on it, to "disregard, override or reverse" a decision, and to stop the system through something like a stop button. In the United States, the voluntary but widely adopted NIST AI Risk Management Framework says much the same in its management function: organizations should have mechanisms "to supersede, disengage, or deactivate" AI systems that start misbehaving.

Most small and mid-size businesses are not deploying high-risk systems in the Act's legal sense, so this is not usually a compliance obligation yet. But it is the direction the standard of care is moving, and it is a fair description of what competent looks like. If you are building something that touches customers, money, or health, designing it so a person can see what it did and stop it is not an advanced feature. It is the baseline.

What "a human in the loop" actually means, and the trap inside it

The phrase gets used loosely, so here is the concrete version. A human in the loop reviews outputs against known-good answers rather than glancing at them, handles the exceptions and edge cases the model fumbles, owns and watches the escalation path so an unsure system routes to a person instead of guessing, monitors for drift as your content and the world change underneath the prompt, tunes the prompts and the surrounding workflow when quality slips, and carries the accountability the last two sections were about. That is a real role with real hours in it. It is also, done well, a smaller and more interesting job than doing the task by hand.

The trap is that a human in the loop can quietly become a human rubber-stamping the loop, and that is a documented failure mode, not a hypothetical. The MIT Sloan Management Review puts it directly: without real insight into how a system reached its answer, "oversight becomes superficial, reducing human involvement to a rubber stamp rather than acting as a critical check." Decades of human-factors research back this up. The classic work on automation bias found that over-reliance on automation "occurs in both naive and expert participants, cannot be prevented by training or instructions," and shows up in teams as well as in individuals. Experience does not inoculate you against it, and neither does a briefing. The only thing that helps is designing the system so the human can genuinely engage: surfacing citations and confidence, flagging the low-certainty cases for closer review, and making the stop and override real rather than theoretical. That is where the guardrails and human-in-the-loop controls on an agentic system earn their place.

How we build it

None of this is an argument against moving work to AI. The productivity numbers are real, and on the right tasks they are large. It is an argument for being honest about where the person still sits. When we scope an AI project, the first questions are the ones this evidence raises: is the task inside the frontier or near its jagged edge, is it high enough volume to earn the build, and is a wrong answer cheap to catch? If it is, we build the review step, the escalation path, and the citations in from the start, and we tell you plainly where a human has to stay in the loop and roughly how much of their time it takes. If it is not, we say so, which is the same honest-scoping stance we bring to everything.

The task can move to the machine. The judgment, the accountability, and the call on when to overrule it stay with a person. Build it that way and AI becomes a real multiplier on your team. Skip that part and you have automated your mistakes and kept the liability. If you are weighing where AI fits in your operation and want a straight read, tell us about the project.

Sources

  1. Stanford HAI, 2025 AI Index Report
  2. U.S. Census Bureau, The Microstructure of AI Diffusion (CES Working Paper 26-25)
  3. U.S. Census Bureau, Large Firms With at Least 20 Employees Biggest AI Users
  4. Dell'Acqua et al., Navigating the Jagged Technological Frontier (Harvard, BCG, Wharton, MIT)
  5. Brynjolfsson, Li & Raymond, Generative AI at Work (NBER Working Paper 31161)
  6. Vectara, Hughes Hallucination Evaluation Model (HHEM) Leaderboard
  7. Stanford RegLab & HAI, Assessing the Reliability of Leading AI Legal Research Tools (arXiv 2405.20362)
  8. Ars Technica, Air Canada must honor refund policy invented by its chatbot (Moffatt v. Air Canada, 2024 BCCRT 149)
  9. Mata v. Avianca, Inc. (S.D.N.Y. 2023), sanctions order
  10. EU AI Act, Article 14 (Human oversight), Regulation (EU) 2024/1689
  11. NIST AI Risk Management Framework (AI RMF 1.0)
  12. MIT Sloan Management Review, How to Avoid Rubber-Stamping AI Recommendations
  13. Parasuraman & Manzey, Complacency and Bias in Human Use of Automation (Human Factors, 2010)

External sources are provided for verification. NavoTech is not affiliated with and does not endorse the organizations cited.


Written by the team at NavoTech Digital Solutions. Have a project or counter-example? Get in touch.