What is human in the loop?
The term has two lives. In machine learning it describes people labelling training data and rating model answers, and that is the sense most definitions online cover. In operations it means something narrower and more useful to an owner: the checkpoint between what an AI system proposes and what actually happens in your ERP, your inbox or your customer's inbox. This guide is about the second meaning.
The vocabulary comes from the European Commission's High-Level Expert Group on AI, whose 2019 ethics guidelines set out three levels of oversight. With a human in the loop, a person can intervene in every decision cycle of the system, which the group itself noted is often neither possible nor desirable. With a human on the loop, a person takes part in designing the system and monitors it while it runs. With a human in command, a person oversees the system as a whole and decides when, and whether, to use it at all.15
Human in the loop vs human on the loop
The difference is timing. With a human in the loop, nothing happens until a person approves it. With a human on the loop, the system acts first and a person monitors the results, samples them and can reverse them. Most companies need both, applied to different actions of the same AI agent.
Take a supplier-invoice agent. Reading the PDF, extracting the fields and matching them to a purchase order can run with a person on the loop, because nothing has left the building and a wrong match is corrected in minutes. Posting the invoice to the ledger or releasing it for payment needs a person in the loop, because a payment to the wrong account is slow and expensive to recover. Oversight is assigned per action, and one agent usually spans two or three tiers.
| Human in the loop | Human on the loop | |
|---|---|---|
| When the person acts | Before the action takes effect | After, by monitoring and sampling |
| What the person sees | Every item, next to its source document | Dashboards, random samples, exceptions |
| Throughput limit | The reviewer's hours | The system's capacity |
| Good for | Payments, customer messages, legal commitments | Classification, internal updates, data checked again downstream |
| Main risk | Approvals turn into rubber stamps under volume | Errors accumulate between checks |
Why AI agents in operations need a person in the loop
Language models sound just as fluent when they are wrong, agents chain several steps so an early mistake travels, and the company is liable for whatever its systems send out. A checkpoint before the costly steps is cheaper than any of those failures.
The companies that build the models say the same. Anthropic's engineering guidance on agents warns that autonomy raises costs and lets errors compound across steps, and recommends checkpoints where the agent pauses for human feedback, plus stopping conditions such as a cap on iterations.8 OpenAI's guide to building agents names two triggers for human intervention: an agent exceeding its failure limits, and actions that are sensitive, irreversible or high-stakes. Its examples are cancelling orders, authorising large refunds and making payments.9
The legal side is less abstract than it sounds. In February 2024 a Canadian tribunal ordered Air Canada to pay a customer CAD 812.02 after its website chatbot gave him wrong information about bereavement fares. The airline argued, in effect, that the chatbot was responsible for its own actions; the tribunal rejected that and held the company responsible for all the information on its website.7 The sum was small. The principle covers every quote, order confirmation and customer reply an agent sends in your name.
Weak risk control is also a common reason projects get cancelled. In June 2025 Gartner predicted that over 40% of agentic AI projects will be cancelled by the end of 2027, and named escalating costs, unclear business value and inadequate risk controls as the causes.10 An approval design written before the build deals with the third cause directly and makes the second measurable.
How to decide which actions need approval
Score each action an agent can take on four questions: can it be undone, how much money does it move or commit, does anyone outside the company see it, and is it a decision about a person. The answers place the action in one of four tiers.
- Undo. Can a mistake be reversed in minutes, at no cost, before anyone notices? A mis-tagged email can. A sent email, an invoice posted to a closed period or a payment cannot.
- Value. What is the largest amount this single action can move or commit? Set a euro limit per action type; above it, the tier goes up.
- Audience. Does the output leave the company, to a customer, supplier, bank, customs office or the tax authority? External output carries your name and, as the Air Canada case shows, your liability.
- People. Does the action decide something about an individual, such as a job applicant, an employee's shifts or a sole trader's credit limit? Then GDPR Article 22 may apply, and the review has to be meaningful.
| Tier | The agent | The person | Typical back-office actions |
|---|---|---|---|
| 1 · Act and log | Acts at once; every action logged | Reviews the log and exceptions weekly | Classifying incoming email, tagging documents, extracting fields into a draft, routing to the right queue |
| 2 · Act, then sample | Acts; a share of items is pulled for review | Checks a random sample daily and can reverse | Creating CRM activities, updating delivery status, matching invoices to purchase orders within tolerance |
| 3 · Draft for approval | Prepares the full action and waits | Approves, edits or rejects each item | Sending quotes, confirming orders, posting supplier invoices, replying to complaints |
| 4 · Human decides | Gathers the facts and drafts a recommendation | Makes the decision; the agent records it | Changing supplier bank details, credit limits, contract terms, anything about an employee or applicant |
Two rules keep the table honest. When the four answers disagree, the strictest one wins: a reversible action that sends money outside the company is tier 3, even if it is easy to undo. And some actions never graduate. A change to a supplier's bank account details is a common route for payment fraud, so it stays with a person however well the agent performs.
The same scoring works outside finance. In a logistics company, reading a CMR consignment note into the transport system is tier 1 or 2, while sending a customer a revised delivery price is tier 3. For an exporter, a quote drafted from the price list and CRM history goes out only after a salesperson approves it.
Setting the review threshold
Do not route items on the model's own confidence score. Route them on checks you can verify: whether the extracted totals add up, whether the VAT number is valid, whether the price matches your price list, whether the customer and the product exist in your system. An item that fails any check goes to a person, whatever the model says about itself.
A language model's self-reported confidence is a weak signal, because the model can be confidently wrong. Deterministic checks are cheap, they explain themselves, and a rejected item arrives with a reason the reviewer can act on. Where a real probability is available, as with a classifier trained on your own labelled history, set the cut-off from data: run a few hundred real items, sort them by score, and choose the point below which the error rate is higher than you accept for that tier.
For a supplier-invoice agent, an illustrative rule set that sends an item to review looks like this:
- line items and VAT do not reconcile to the invoice total;
- the VAT number fails validation, or the supplier is not in your master data;
- the bank account on the invoice differs from the one on file;
- there is no matching purchase order, or the price differs beyond the agreed tolerance;
- the amount is above the limit set for automatic processing;
- the invoice number already exists for that supplier.
The agent also needs a way to decline. The abstention path, where the agent stops, states which check failed and hands the item over, is the most important branch in the design. The company world model described in our AI ontology long-read works the same way: when context is missing or an action is risky under your rules, the correct output is no action.
How to stop approvals turning into rubber-stamping
People who review machine output learn to trust it, and then stop checking. The approval screen has to be designed for attention: show the evidence, flag only what failed, make each approval a deliberate act, and measure whether reviewers catch errors.
This is automation bias, and it is well documented. A 2010 review in Human Factors found it in novices and experts alike and concluded that training or instructions alone do not prevent it.11 In a 2023 study in Radiology, when an AI system suggested the wrong assessment for a mammogram, the accuracy of inexperienced radiologists fell from 79.7% to 19.8%, and that of very experienced ones from 82.3% to 45.5%.13 A meta-analysis of clinical decision-support studies found that wrong advice raised the risk of a wrong decision by 26%, and that stressing the reviewer's accountability and showing the system's confidence reduced the effect.12 Those studies are medical, but the mechanism is the same one an accounts clerk faces at 16:45 with forty invoices still in the queue.
Standards bodies and regulators have noticed. NIST's risk profile for generative AI lists automation bias and over-reliance among its human-AI configuration risks,14 the EU AI Act names automation bias in its human-oversight article,1 and European data-protection guidance says human involvement must be "meaningful, rather than just a token gesture", carried out by someone with the authority and competence to change the decision.5 Design rules that follow from this:
- Put the source next to the proposal. The reviewer sees the PDF or email beside the extracted fields, never the fields alone.
- Flag what failed. Highlight the two fields that broke a check instead of asking for a review of all thirty.
- No bulk approve in tier 3. Each item needs its own action, and the queue a person clears in one sitting is capped.
- Capture a reason. Every edit or rejection takes a short reason code; those codes become your error statistics.
- Test the tester. Now and then, insert a known-bad item with a planted error and record whether it is caught.
- Log who approved what, and how long it took. A two-second approval of a thirty-field invoice is a signal worth reviewing.
When a step can move to less oversight
Only on measured evidence agreed in advance: a sample size, an error limit, a person who signs off the change, and a rule that sends the step back if errors return. Loosening oversight because the agent "has been fine lately" is how silent errors reach customers.
A simple statistical rule helps with the sample size. If you check n consecutive cases and find no errors, the upper end of the 95% confidence interval for the true error rate is roughly 3 ÷ n.16 Fifty clean cases therefore show only that the error rate is probably below 6%. To claim below 2% you need about 150 clean cases, and to claim below 1%, about 300. Volume decides how long that takes: a step handling 20 items a working day reaches 300 in three weeks.
| Stage | The agent | The people | Move on when |
|---|---|---|---|
| Shadow | Proposes; nothing takes effect | Work as before; outputs are compared | Proposals match the human result at the agreed rate on a labelled sample |
| Approval | Prepares every action | Approve or edit each item | Edit rate stays under the limit over the agreed number of consecutive items |
| Sampled | Acts | Review a random share daily | Sampled error rate stays under the limit, with no errors seen outside the company |
| Exceptions only | Acts; routes failed checks | Handle exceptions, audit monthly | Stays here while checks and monthly audits pass |
This is the rollout Vertara uses on client builds: shadow mode first, then human approval, then supervised autonomy, one action class at a time (the AI agency in Lithuania page describes the full engagement). Two rules sit alongside it. A tier 3 error that reaches a customer, supplier or authority sends that action back to approval immediately. And a change in the inputs, such as a new supplier layout, a new price list or an ERP upgrade, restarts the count for the affected step.
What the EU AI Act and GDPR require
The EU AI Act's human-oversight rules bind high-risk AI systems, and those obligations now start on 2 December 2027. Most back-office agents that read invoices, enter orders or draft quotes are not high-risk. GDPR is the law more likely to apply today, whenever an automated decision significantly affects a person.
EU AI Act, Articles 14 and 26
Article 14 requires high-risk AI systems to be designed so that people can oversee them effectively. The people assigned must be able to understand the system's capacities and limits, stay aware of automation bias, interpret its output correctly, decide to disregard, override or reverse that output, and stop the system.1 Article 26 puts the matching duty on deployers, meaning the companies that use the system: oversight goes to people with the competence, training and authority to exercise it.1 High-risk means the uses listed in Annex III of the Act, such as recruitment and worker management or the creditworthiness of individuals, plus AI built into products covered by EU product-safety law.
The dates moved this summer. Regulation (EU) 2026/1744, in force since 27 July 2026, pushed the start of the high-risk obligations from 2 August 2026 to 2 December 2027 for Annex III systems, and to 2 August 2028 for AI in regulated products.23 The same regulation softened the AI-literacy duty in Article 4, which has applied to every provider and deployer since 2 February 2025: companies must now take measures to build AI literacy among their staff rather than guarantee a set level of it.2
Where Article 14 does not bind you, its five capabilities still make a good checklist for any approval design, and they cost little to build in from the start.
GDPR Article 22
Article 22 gives people the right not to be subject to a decision based solely on automated processing that produces legal or similarly significant effects for them.4 A human reviewer takes a decision out of that category only if the review is real: carried out by someone with the authority to change the outcome, who considers the relevant data.5 In the 2023 SCHUFA judgment, the EU Court of Justice held that even an automated credit score can count as such a decision when the lender relies heavily on it.6 For back-office agents the test is usually simple. Invoices and orders between companies are rarely decisions about people; screening job applicants, setting a sole trader's credit terms or scheduling employees can be.
In Lithuania
The Communications Regulatory Authority (RRT) was designated as the AI Act's national market surveillance authority and single point of contact from 1 April 2025, with the Innovation Agency as the notifying authority; a draft implementing law would spread market surveillance across sector regulators, with RRT remaining the contact point.17 The State Data Protection Inspectorate (VDAI) supervises GDPR and has published guidance in Lithuanian on AI and the right to human intervention.18
How to add human in the loop to an AI workflow
This is the order we follow when we design an agent's approval layer. It happens before any build starts, and it is where most of the risk gets removed.
- List the actions, not the system
Write down every action the agent could take: read, extract, draft, write to a system, send, pay. Each one is scored separately.
- Score and tier each action
Apply the four questions (undo, value, audience, people). The strictest answer sets the tier; mark the actions that never graduate.
- Write the checks and the abstain path
For each action, list the validations that must pass before it proceeds, and what the agent does when one fails: stop, state the reason, hand over.
- Run shadow mode on real work
The agent processes live items but nothing takes effect. Compare its output with what your staff did, field by field, and label the differences.
- Design the approval screen for attention
Source next to proposal, failed fields flagged, one action per item, and a reason code on every edit or rejection.
- Agree graduation and rollback in writing
Sample size, error limit per tier, who signs off a change of tier, and the events that send a step back to approval.
- Keep sampling after go-live
Even at the lightest tier, pull a random sample every month and restart the count when inputs change. If you want this designed and built into your own systems, see business process automation with Vertara.
Common questions.
What is human in the loop in AI?
Human in the loop is a way of running an AI system where a person reviews, corrects or approves its output before it takes effect. In business operations, an AI agent prepares an invoice posting, order or reply, and a named employee signs it off. In machine learning, the same term describes people labelling training data.
What is the difference between human in the loop and human on the loop?
Timing. With a human in the loop, nothing takes effect until a person approves it. With a human on the loop, the system acts first and a person monitors, samples and can reverse its work. Use the first for actions that are hard to undo or leave the company, and the second for high-volume internal steps.
What is an example of human in the loop in a business?
A supplier-invoice agent reads the PDF, extracts the fields and matches them to the purchase order on its own, then prepares the ledger posting and waits. An accounts clerk sees the invoice next to the proposed posting, checks the flagged fields and approves or corrects it. Only the approved posting reaches the accounting system.
Does the EU AI Act require human in the loop?
Article 14 requires human oversight for high-risk AI systems, such as those used in recruitment or credit scoring, and those obligations now start on 2 December 2027. Most back-office agents for invoices, orders or quotes are not high-risk. GDPR Article 22 separately requires meaningful human review of automated decisions with significant effects on people.
What confidence threshold should trigger human review?
Avoid relying on a language model's self-reported confidence. Send an item to review whenever a verifiable check fails: the totals do not reconcile, the VAT number is invalid, the price differs from the price list, the bank account has changed or the amount exceeds a set limit. Where a calibrated score exists, set the cut-off from a labelled sample of your own cases.
How do you stop reviewers from rubber-stamping AI output?
Show the source document next to the proposal, flag only the fields that failed a check, require a separate action per item instead of bulk approval, capture a reason for every edit, occasionally insert a known-bad test item, and log how long each approval takes. Research on automation bias shows that training alone does not remove it.
Is human in the loop the same as RLHF?
No. Reinforcement learning from human feedback (RLHF) is a training method: people rate model answers and the ratings are used to tune the model. Human in the loop in operations is a runtime control: a person approves what a deployed system is about to do. A company using an AI agent needs the second, and the model vendor handles the first.
References
- European Union - Regulation (EU) 2024/1689 (Artificial Intelligence Act), Articles 4, 14 and 26 and Annex III, OJ L, 12 July 2024. eur-lex.europa.eu/eli/reg/2024/1689/oj
- European Union - Regulation (EU) 2026/1744 amending Regulation (EU) 2024/1689 (Digital Omnibus on AI), OJ L, 24 July 2026. eur-lex.europa.eu/eli/reg/2026/1744/oj
- European Commission - "AI Omnibus enters into force", July 2026. digital-strategy.ec.europa.eu/en/news/ai-omnibus-enters-force
- European Union - Regulation (EU) 2016/679 (General Data Protection Regulation), Article 22, OJ L 119, 4 May 2016. eur-lex.europa.eu/eli/reg/2016/679/oj
- Article 29 Working Party - "Guidelines on Automated individual decision-making and Profiling" (WP251rev.01), revised 6 February 2018, endorsed by the EDPB on 25 May 2018, p. 21. ec.europa.eu/newsroom/article29/redirection/document/49826
- Court of Justice of the EU - Judgment in Case C-634/21 (SCHUFA Holding, scoring), 7 December 2023; Press Release 186/23. curia.europa.eu/…/cp230186en.pdf
- BC Civil Resolution Tribunal - Moffatt v. Air Canada, 2024 BCCRT 149, 14 February 2024, paras 27 and 44. decisions.civilresolutionbc.ca/crt/crtd/en/item/525448
- Anthropic - "Building effective agents", 19 December 2024. anthropic.com/engineering/building-effective-agents
- OpenAI - "A practical guide to building agents", 2025, sections "Guardrails" and "Plan for human intervention". cdn.openai.com/…/a-practical-guide-to-building-agents.pdf
- Gartner - "Gartner Predicts Over 40% of Agentic AI Projects Will Be Canceled by End of 2027", press release, 25 June 2025. gartner.com/en/newsroom/press-releases/2025-06-25-…
- Parasuraman, R. and Manzey, D. H. - "Complacency and Bias in Human Use of Automation: An Attentional Integration", Human Factors 52(3), 381–410, June 2010. doi.org/10.1177/0018720810376055
- Goddard, K., Roudsari, A. and Wyatt, J. C. - "Automation bias: a systematic review of frequency, effect mediators, and mitigators", JAMIA 19(1), 121–127, 2012. pmc.ncbi.nlm.nih.gov/articles/PMC3240751
- Dratsch, T. et al. - "Automation Bias in Mammography: The Impact of Artificial Intelligence BI-RADS Suggestions on Reader Performance", Radiology 307(4), 2 May 2023. pubmed.ncbi.nlm.nih.gov/37129490
- NIST - "Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile" (NIST AI 600-1), July 2024, risk 7 "Human-AI Configuration". nvlpubs.nist.gov/nistpubs/ai/NIST.AI.600-1.pdf
- High-Level Expert Group on AI (European Commission) - "Ethics Guidelines for Trustworthy AI", 8 April 2019, section on human agency and oversight. digital-strategy.ec.europa.eu/en/library/ethics-guidelines-trustworthy-ai
- Hanley, J. A. and Lippman-Hand, A. - "If nothing goes wrong, is everything all right? Interpreting zero numerators", JAMA 249(13), 1743–1745, 1 April 1983. pubmed.ncbi.nlm.nih.gov/6827763
- Ryšių reguliavimo tarnyba (RRT) - "Ryšių reguliavimo tarnyba taps pagrindine dirbtinio intelekto priežiūros institucija Lietuvoje", 16 January 2025 (archived copy), and "DI reguliavimas" (in Lithuanian). rrt.lt/veiklos-sritys/skaitmenine-erdve/di-informacija/di-reguliavimas
- Valstybinė duomenų apsaugos inspekcija (VDAI) - "DUK. Dirbtinis intelektas" (FAQ on AI and personal data, in Lithuanian), 29 May 2025. vdai.lrv.lt/…/2025-05-29 DUK del DI sprendimu