Home/Insights/Human in the loop

Human in the loop: how to design approval tiers for AI

Human in the loop means a person reviews or approves an AI system's output before it takes effect. The hard part is deciding where. This guide sorts back-office actions into four approval tiers by reversibility, money and who sees the result, shows how to set review thresholds from measured errors, and explains what the EU AI Act and GDPR actually require.

A conveyor from the Vertara object family carrying white boxes between two terrazzo gates, each box passing a mint checkpoint
Definition

What is human in the loop?

Human in the loop (HITL) is a way of running an AI system in which a person reviews, corrects or approves its output before that output takes effect. In business operations, the AI agent prepares the invoice posting, order or customer reply, and a named employee signs it off.

The term has two lives. In machine learning it describes people labelling training data and rating model answers, and that is the sense most definitions online cover. In operations it means something narrower and more useful to an owner: the checkpoint between what an AI system proposes and what actually happens in your ERP, your inbox or your customer's inbox. This guide is about the second meaning.

The vocabulary comes from the European Commission's High-Level Expert Group on AI, whose 2019 ethics guidelines set out three levels of oversight. With a human in the loop, a person can intervene in every decision cycle of the system, which the group itself noted is often neither possible nor desirable. With a human on the loop, a person takes part in designing the system and monitors it while it runs. With a human in command, a person oversees the system as a whole and decides when, and whether, to use it at all.15

Human in the loopapprove before it happens
The AI proposes; a person approves, edits or rejects each item before it takes effect. Right for actions that are hard to undo or leave the company.
Human on the loopmonitor after
The AI acts; a person watches, samples and can stop or reverse. Right for high-volume actions that are cheap to correct.
Human in commanddecide whether to use it
A person decides which tasks the system may touch at all, and can switch it off.
Shadow modeprove before trusting
The AI runs on real work, but its output is only compared with what staff did. Nothing it produces takes effect.
Automation biasthe failure mode
The tendency to accept a machine's suggestion without checking it, including when it is wrong. The EU AI Act names it in Article 14.
Abstentionthe safe default
What the agent does when a check fails: it stops, says which check failed and hands the item to a person.
The distinction

Human in the loop vs human on the loop

The difference is timing. With a human in the loop, nothing happens until a person approves it. With a human on the loop, the system acts first and a person monitors the results, samples them and can reverse them. Most companies need both, applied to different actions of the same AI agent.

Take a supplier-invoice agent. Reading the PDF, extracting the fields and matching them to a purchase order can run with a person on the loop, because nothing has left the building and a wrong match is corrected in minutes. Posting the invoice to the ledger or releasing it for payment needs a person in the loop, because a payment to the wrong account is slow and expensive to recover. Oversight is assigned per action, and one agent usually spans two or three tiers.

Human in the loopHuman on the loop
When the person actsBefore the action takes effectAfter, by monitoring and sampling
What the person seesEvery item, next to its source documentDashboards, random samples, exceptions
Throughput limitThe reviewer's hoursThe system's capacity
Good forPayments, customer messages, legal commitmentsClassification, internal updates, data checked again downstream
Main riskApprovals turn into rubber stamps under volumeErrors accumulate between checks
The case for it

Why AI agents in operations need a person in the loop

Language models sound just as fluent when they are wrong, agents chain several steps so an early mistake travels, and the company is liable for whatever its systems send out. A checkpoint before the costly steps is cheaper than any of those failures.

The companies that build the models say the same. Anthropic's engineering guidance on agents warns that autonomy raises costs and lets errors compound across steps, and recommends checkpoints where the agent pauses for human feedback, plus stopping conditions such as a cap on iterations.8 OpenAI's guide to building agents names two triggers for human intervention: an agent exceeding its failure limits, and actions that are sensitive, irreversible or high-stakes. Its examples are cancelling orders, authorising large refunds and making payments.9

The legal side is less abstract than it sounds. In February 2024 a Canadian tribunal ordered Air Canada to pay a customer CAD 812.02 after its website chatbot gave him wrong information about bereavement fares. The airline argued, in effect, that the chatbot was responsible for its own actions; the tribunal rejected that and held the company responsible for all the information on its website.7 The sum was small. The principle covers every quote, order confirmation and customer reply an agent sends in your name.

Weak risk control is also a common reason projects get cancelled. In June 2025 Gartner predicted that over 40% of agentic AI projects will be cancelled by the end of 2027, and named escalating costs, unclear business value and inadequate risk controls as the causes.10 An approval design written before the build deals with the third cause directly and makes the second measurable.

40%+
of agentic AI projects Gartner expects to be cancelled by the end of 2027
Gartner · Jun 2025
CAD 812
awarded against Air Canada for its chatbot's wrong answer; the company was liable, not the bot
BC Civil Resolution Tribunal · Feb 2024
2 Dec 2027
new start date for the EU AI Act's high-risk obligations, including human oversight
Regulation (EU) 2026/1744
The tiers

How to decide which actions need approval

Score each action an agent can take on four questions: can it be undone, how much money does it move or commit, does anyone outside the company see it, and is it a decision about a person. The answers place the action in one of four tiers.

  • Undo. Can a mistake be reversed in minutes, at no cost, before anyone notices? A mis-tagged email can. A sent email, an invoice posted to a closed period or a payment cannot.
  • Value. What is the largest amount this single action can move or commit? Set a euro limit per action type; above it, the tier goes up.
  • Audience. Does the output leave the company, to a customer, supplier, bank, customs office or the tax authority? External output carries your name and, as the Air Canada case shows, your liability.
  • People. Does the action decide something about an individual, such as a job applicant, an employee's shifts or a sole trader's credit limit? Then GDPR Article 22 may apply, and the review has to be meaningful.
Four approval tiers for back-office AI
EASY TO UNDO · INTERNAL · SMALL AMOUNTS HARD TO UNDO · EXTERNAL · LARGE · ABOUT PEOPLE TIER 1 ACT AND LOG classify email · tag docs extract fields to a draft TIER 2 ACT, THEN SAMPLE CRM activities delivery status updates invoice-to-PO match within tolerance person on the loop TIER 3 DRAFT FOR APPROVAL send quotes confirm orders post supplier invoices reply to complaints person in the loop TIER 4 HUMAN DECIDES supplier bank details credit limits contract terms anything about an employee or applicant AI drafts, person decides
Oversight rises with the cost of a mistake. Each action an agent can take is placed separately, so one invoice agent can run tier 1 for extraction and tier 3 for posting.
TierThe agentThe personTypical back-office actions
1 · Act and logActs at once; every action loggedReviews the log and exceptions weeklyClassifying incoming email, tagging documents, extracting fields into a draft, routing to the right queue
2 · Act, then sampleActs; a share of items is pulled for reviewChecks a random sample daily and can reverseCreating CRM activities, updating delivery status, matching invoices to purchase orders within tolerance
3 · Draft for approvalPrepares the full action and waitsApproves, edits or rejects each itemSending quotes, confirming orders, posting supplier invoices, replying to complaints
4 · Human decidesGathers the facts and drafts a recommendationMakes the decision; the agent records itChanging supplier bank details, credit limits, contract terms, anything about an employee or applicant

Two rules keep the table honest. When the four answers disagree, the strictest one wins: a reversible action that sends money outside the company is tier 3, even if it is easy to undo. And some actions never graduate. A change to a supplier's bank account details is a common route for payment fraud, so it stays with a person however well the agent performs.

The same scoring works outside finance. In a logistics company, reading a CMR consignment note into the transport system is tier 1 or 2, while sending a customer a revised delivery price is tier 3. For an exporter, a quote drafted from the price list and CRM history goes out only after a salesperson approves it.

The threshold

Setting the review threshold

Do not route items on the model's own confidence score. Route them on checks you can verify: whether the extracted totals add up, whether the VAT number is valid, whether the price matches your price list, whether the customer and the product exist in your system. An item that fails any check goes to a person, whatever the model says about itself.

A language model's self-reported confidence is a weak signal, because the model can be confidently wrong. Deterministic checks are cheap, they explain themselves, and a rejected item arrives with a reason the reviewer can act on. Where a real probability is available, as with a classifier trained on your own labelled history, set the cut-off from data: run a few hundred real items, sort them by score, and choose the point below which the error rate is higher than you accept for that tier.

For a supplier-invoice agent, an illustrative rule set that sends an item to review looks like this:

  • line items and VAT do not reconcile to the invoice total;
  • the VAT number fails validation, or the supplier is not in your master data;
  • the bank account on the invoice differs from the one on file;
  • there is no matching purchase order, or the price differs beyond the agreed tolerance;
  • the amount is above the limit set for automatic processing;
  • the invoice number already exists for that supplier.

The agent also needs a way to decline. The abstention path, where the agent stops, states which check failed and hands the item over, is the most important branch in the design. The company world model described in our AI ontology long-read works the same way: when context is missing or an action is risky under your rules, the correct output is no action.

The reviewer

How to stop approvals turning into rubber-stamping

People who review machine output learn to trust it, and then stop checking. The approval screen has to be designed for attention: show the evidence, flag only what failed, make each approval a deliberate act, and measure whether reviewers catch errors.

This is automation bias, and it is well documented. A 2010 review in Human Factors found it in novices and experts alike and concluded that training or instructions alone do not prevent it.11 In a 2023 study in Radiology, when an AI system suggested the wrong assessment for a mammogram, the accuracy of inexperienced radiologists fell from 79.7% to 19.8%, and that of very experienced ones from 82.3% to 45.5%.13 A meta-analysis of clinical decision-support studies found that wrong advice raised the risk of a wrong decision by 26%, and that stressing the reviewer's accountability and showing the system's confidence reduced the effect.12 Those studies are medical, but the mechanism is the same one an accounts clerk faces at 16:45 with forty invoices still in the queue.

Standards bodies and regulators have noticed. NIST's risk profile for generative AI lists automation bias and over-reliance among its human-AI configuration risks,14 the EU AI Act names automation bias in its human-oversight article,1 and European data-protection guidance says human involvement must be "meaningful, rather than just a token gesture", carried out by someone with the authority and competence to change the decision.5 Design rules that follow from this:

  • Put the source next to the proposal. The reviewer sees the PDF or email beside the extracted fields, never the fields alone.
  • Flag what failed. Highlight the two fields that broke a check instead of asking for a review of all thirty.
  • No bulk approve in tier 3. Each item needs its own action, and the queue a person clears in one sitting is capped.
  • Capture a reason. Every edit or rejection takes a short reason code; those codes become your error statistics.
  • Test the tester. Now and then, insert a known-bad item with a planted error and record whether it is caught.
  • Log who approved what, and how long it took. A two-second approval of a thirty-field invoice is a signal worth reviewing.
Graduation

When a step can move to less oversight

Only on measured evidence agreed in advance: a sample size, an error limit, a person who signs off the change, and a rule that sends the step back if errors return. Loosening oversight because the agent "has been fine lately" is how silent errors reach customers.

A simple statistical rule helps with the sample size. If you check n consecutive cases and find no errors, the upper end of the 95% confidence interval for the true error rate is roughly 3 ÷ n.16 Fifty clean cases therefore show only that the error rate is probably below 6%. To claim below 2% you need about 150 clean cases, and to claim below 1%, about 300. Volume decides how long that takes: a step handling 20 items a working day reaches 300 in three weeks.

StageThe agentThe peopleMove on when
ShadowProposes; nothing takes effectWork as before; outputs are comparedProposals match the human result at the agreed rate on a labelled sample
ApprovalPrepares every actionApprove or edit each itemEdit rate stays under the limit over the agreed number of consecutive items
SampledActsReview a random share dailySampled error rate stays under the limit, with no errors seen outside the company
Exceptions onlyActs; routes failed checksHandle exceptions, audit monthlyStays here while checks and monthly audits pass

This is the rollout Vertara uses on client builds: shadow mode first, then human approval, then supervised autonomy, one action class at a time (the AI agency in Lithuania page describes the full engagement). Two rules sit alongside it. A tier 3 error that reaches a customer, supplier or authority sends that action back to approval immediately. And a change in the inputs, such as a new supplier layout, a new price list or an ERP upgrade, restarts the count for the affected step.

The law

What the EU AI Act and GDPR require

The EU AI Act's human-oversight rules bind high-risk AI systems, and those obligations now start on 2 December 2027. Most back-office agents that read invoices, enter orders or draft quotes are not high-risk. GDPR is the law more likely to apply today, whenever an automated decision significantly affects a person.

EU AI Act, Articles 14 and 26

Article 14 requires high-risk AI systems to be designed so that people can oversee them effectively. The people assigned must be able to understand the system's capacities and limits, stay aware of automation bias, interpret its output correctly, decide to disregard, override or reverse that output, and stop the system.1 Article 26 puts the matching duty on deployers, meaning the companies that use the system: oversight goes to people with the competence, training and authority to exercise it.1 High-risk means the uses listed in Annex III of the Act, such as recruitment and worker management or the creditworthiness of individuals, plus AI built into products covered by EU product-safety law.

The dates moved this summer. Regulation (EU) 2026/1744, in force since 27 July 2026, pushed the start of the high-risk obligations from 2 August 2026 to 2 December 2027 for Annex III systems, and to 2 August 2028 for AI in regulated products.23 The same regulation softened the AI-literacy duty in Article 4, which has applied to every provider and deployer since 2 February 2025: companies must now take measures to build AI literacy among their staff rather than guarantee a set level of it.2

Where Article 14 does not bind you, its five capabilities still make a good checklist for any approval design, and they cost little to build in from the start.

GDPR Article 22

Article 22 gives people the right not to be subject to a decision based solely on automated processing that produces legal or similarly significant effects for them.4 A human reviewer takes a decision out of that category only if the review is real: carried out by someone with the authority to change the outcome, who considers the relevant data.5 In the 2023 SCHUFA judgment, the EU Court of Justice held that even an automated credit score can count as such a decision when the lender relies heavily on it.6 For back-office agents the test is usually simple. Invoices and orders between companies are rarely decisions about people; screening job applicants, setting a sole trader's credit terms or scheduling employees can be.

In Lithuania

The Communications Regulatory Authority (RRT) was designated as the AI Act's national market surveillance authority and single point of contact from 1 April 2025, with the Innovation Agency as the notifying authority; a draft implementing law would spread market surveillance across sector regulators, with RRT remaining the contact point.17 The State Data Protection Inspectorate (VDAI) supervises GDPR and has published guidance in Lithuanian on AI and the right to human intervention.18

The playbook

How to add human in the loop to an AI workflow

This is the order we follow when we design an agent's approval layer. It happens before any build starts, and it is where most of the risk gets removed.

  1. List the actions, not the system

    Write down every action the agent could take: read, extract, draft, write to a system, send, pay. Each one is scored separately.

  2. Score and tier each action

    Apply the four questions (undo, value, audience, people). The strictest answer sets the tier; mark the actions that never graduate.

  3. Write the checks and the abstain path

    For each action, list the validations that must pass before it proceeds, and what the agent does when one fails: stop, state the reason, hand over.

  4. Run shadow mode on real work

    The agent processes live items but nothing takes effect. Compare its output with what your staff did, field by field, and label the differences.

  5. Design the approval screen for attention

    Source next to proposal, failed fields flagged, one action per item, and a reason code on every edit or rejection.

  6. Agree graduation and rollback in writing

    Sample size, error limit per tier, who signs off a change of tier, and the events that send a step back to approval.

  7. Keep sampling after go-live

    Even at the lightest tier, pull a random sample every month and restart the count when inputs change. If you want this designed and built into your own systems, see business process automation with Vertara.

Common questions.

What is human in the loop in AI?

Human in the loop is a way of running an AI system where a person reviews, corrects or approves its output before it takes effect. In business operations, an AI agent prepares an invoice posting, order or reply, and a named employee signs it off. In machine learning, the same term describes people labelling training data.

What is the difference between human in the loop and human on the loop?

Timing. With a human in the loop, nothing takes effect until a person approves it. With a human on the loop, the system acts first and a person monitors, samples and can reverse its work. Use the first for actions that are hard to undo or leave the company, and the second for high-volume internal steps.

What is an example of human in the loop in a business?

A supplier-invoice agent reads the PDF, extracts the fields and matches them to the purchase order on its own, then prepares the ledger posting and waits. An accounts clerk sees the invoice next to the proposed posting, checks the flagged fields and approves or corrects it. Only the approved posting reaches the accounting system.

Does the EU AI Act require human in the loop?

Article 14 requires human oversight for high-risk AI systems, such as those used in recruitment or credit scoring, and those obligations now start on 2 December 2027. Most back-office agents for invoices, orders or quotes are not high-risk. GDPR Article 22 separately requires meaningful human review of automated decisions with significant effects on people.

What confidence threshold should trigger human review?

Avoid relying on a language model's self-reported confidence. Send an item to review whenever a verifiable check fails: the totals do not reconcile, the VAT number is invalid, the price differs from the price list, the bank account has changed or the amount exceeds a set limit. Where a calibrated score exists, set the cut-off from a labelled sample of your own cases.

How do you stop reviewers from rubber-stamping AI output?

Show the source document next to the proposal, flag only the fields that failed a check, require a separate action per item instead of bulk approval, capture a reason for every edit, occasionally insert a known-bad test item, and log how long each approval takes. Research on automation bias shows that training alone does not remove it.

Is human in the loop the same as RLHF?

No. Reinforcement learning from human feedback (RLHF) is a training method: people rate model answers and the ratings are used to tune the model. Human in the loop in operations is a runtime control: a person approves what a deployed system is about to do. A company using an AI agent needs the second, and the model vendor handles the first.

Aurelijus Jakas
Aurelijus Jakas
Founder · Vertara

Designs and builds custom AI systems for mid-sized manufacturing, logistics, export and professional-services companies in Lithuania and across Europe. Researched with AI tools; written and checked by the author. All articles

Appendix · Sources

References

  1. European Union - Regulation (EU) 2024/1689 (Artificial Intelligence Act), Articles 4, 14 and 26 and Annex III, OJ L, 12 July 2024. eur-lex.europa.eu/eli/reg/2024/1689/oj
  2. European Union - Regulation (EU) 2026/1744 amending Regulation (EU) 2024/1689 (Digital Omnibus on AI), OJ L, 24 July 2026. eur-lex.europa.eu/eli/reg/2026/1744/oj
  3. European Commission - "AI Omnibus enters into force", July 2026. digital-strategy.ec.europa.eu/en/news/ai-omnibus-enters-force
  4. European Union - Regulation (EU) 2016/679 (General Data Protection Regulation), Article 22, OJ L 119, 4 May 2016. eur-lex.europa.eu/eli/reg/2016/679/oj
  5. Article 29 Working Party - "Guidelines on Automated individual decision-making and Profiling" (WP251rev.01), revised 6 February 2018, endorsed by the EDPB on 25 May 2018, p. 21. ec.europa.eu/newsroom/article29/redirection/document/49826
  6. Court of Justice of the EU - Judgment in Case C-634/21 (SCHUFA Holding, scoring), 7 December 2023; Press Release 186/23. curia.europa.eu/…/cp230186en.pdf
  7. BC Civil Resolution Tribunal - Moffatt v. Air Canada, 2024 BCCRT 149, 14 February 2024, paras 27 and 44. decisions.civilresolutionbc.ca/crt/crtd/en/item/525448
  8. Anthropic - "Building effective agents", 19 December 2024. anthropic.com/engineering/building-effective-agents
  9. OpenAI - "A practical guide to building agents", 2025, sections "Guardrails" and "Plan for human intervention". cdn.openai.com/…/a-practical-guide-to-building-agents.pdf
  10. Gartner - "Gartner Predicts Over 40% of Agentic AI Projects Will Be Canceled by End of 2027", press release, 25 June 2025. gartner.com/en/newsroom/press-releases/2025-06-25-…
  11. Parasuraman, R. and Manzey, D. H. - "Complacency and Bias in Human Use of Automation: An Attentional Integration", Human Factors 52(3), 381–410, June 2010. doi.org/10.1177/0018720810376055
  12. Goddard, K., Roudsari, A. and Wyatt, J. C. - "Automation bias: a systematic review of frequency, effect mediators, and mitigators", JAMIA 19(1), 121–127, 2012. pmc.ncbi.nlm.nih.gov/articles/PMC3240751
  13. Dratsch, T. et al. - "Automation Bias in Mammography: The Impact of Artificial Intelligence BI-RADS Suggestions on Reader Performance", Radiology 307(4), 2 May 2023. pubmed.ncbi.nlm.nih.gov/37129490
  14. NIST - "Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile" (NIST AI 600-1), July 2024, risk 7 "Human-AI Configuration". nvlpubs.nist.gov/nistpubs/ai/NIST.AI.600-1.pdf
  15. High-Level Expert Group on AI (European Commission) - "Ethics Guidelines for Trustworthy AI", 8 April 2019, section on human agency and oversight. digital-strategy.ec.europa.eu/en/library/ethics-guidelines-trustworthy-ai
  16. Hanley, J. A. and Lippman-Hand, A. - "If nothing goes wrong, is everything all right? Interpreting zero numerators", JAMA 249(13), 1743–1745, 1 April 1983. pubmed.ncbi.nlm.nih.gov/6827763
  17. Ryšių reguliavimo tarnyba (RRT) - "Ryšių reguliavimo tarnyba taps pagrindine dirbtinio intelekto priežiūros institucija Lietuvoje", 16 January 2025 (archived copy), and "DI reguliavimas" (in Lithuanian). rrt.lt/veiklos-sritys/skaitmenine-erdve/di-informacija/di-reguliavimas
  18. Valstybinė duomenų apsaugos inspekcija (VDAI) - "DUK. Dirbtinis intelektas" (FAQ on AI and personal data, in Lithuanian), 29 May 2025. vdai.lrv.lt/…/2025-05-29 DUK del DI sprendimu

Show us the workflow. We'll tier it.

A free 30-minute call. Bring one process: we'll tell you which steps an agent can run alone, which need a person's approval, and whether the build clears the 60% rule.

  • A concrete read on your highest-ROI opportunity
  • A ballpark of the investment and the return
  • A clear next step, only if it makes sense

Prefer email? [email protected]

Tell us about your workflow

You're not committing to a build. On the call we tell you straight whether the math clears — and we don't sell a build that won't pay back.

We reply within one business day · Your details stay private