Say the AI on your desk writes a reply to a customer at 9:02 on a Monday. Human in the loop is simply whether one of your people reads that reply before the customer does. On most tools the answer hangs on a confidence threshold, a bar the AI’s own sureness has to clear before it is let send anything. Replies over the bar go out. The rest wait in the ticket as drafts for a person, and where the bar sits decides which mistakes your desk makes a good deal more than how many.
Setting it feels like a question of nerve in the first week, how brave are we feeling this morning. It is really a counting job, truth be told, and the counting most teams do in that first month is the wrong way round, which is why their bar only ever seems to go up.
What does a confidence threshold actually decide
Here is a Tuesday on a small shop’s desk, with two things going wrong on it before lunch.
At 8:40 somebody asks whether they can still swap a pair of boots for the next size up. The AI is sure of itself, and inside 4 minutes a reply has gone out on its own, polite and well put together, telling them yes. The reply was lovely. The boots were bought in a sale that did not allow exchanges. Nobody on your team saw it go. The customer finds out when the parcel comes back to them with a note in it.
At 9:15 somebody else asks where their order is. This time the AI writes a perfectly good answer with the tracking link in it and then sits on it, its confidence having come in a hair under the bar. Whoever is on the desk gets round to it after lunch, reads it through, changes not a word, and goes and sends it. The customer had been refreshing Gmail on the bus since a quarter past nine. Nobody ever writes in about that sort of wait, so nobody ever counts it, and it never once turns up in any of the reports the desk goes and looks at on a Monday morning.
Put the bar up after the boots and you will get fewer boots, mind you. You will also get a good many more tracking links sitting in tickets till after lunch, and on the made-up shop further down that came to 74 of them in a month. Bring it back down and the boots come back. No setting we know of gets rid of both.
The two do not cost the same, and that is the heart of it. The tracking link cost the customer 4 hours, which shows up on your first response time and against the SLA, and maybe leaves them a shade cooler with you than they were at breakfast. As for the boots, there is a return label to pay for, a pair that will not sell again at full price, and a customer who reads your next email with one eyebrow up. Get a refund answer wrong and it can cost you the refund its self. So one bar for the whole desk will be wrong somewhere, on account of a mistake costing next to nothing on one sort of question and real money on another.
Why a held draft sent unchanged is the number to watch
In the first month nearly every team keeps its eye on the wrong pile. It is an easy one to go and make. The held drafts are the visible ones, somebody has to deal with each of them, and every time a person edits a draft before sending it, it feels like proof the bar is doing its job.
Sit behind somebody clearing held drafts for 20 minutes and watch what the edits actually are. The first one softens a greeting. On the next, the customer’s first name gets pulled across from the CRM, and a few drafts later “we apologise for any inconvenience” comes out and “sorry about that” goes in, funnily enough, before it is sent. Nothing a customer could act on changed in any of them. Had the AI sent those on its own, the customer would have got the very same answer in slightly different clothes and not minded one bit.
What we would do with a month of them is go through with a pen and mark every held draft where a person changed a fact, whether a date, an amount, or a yes that ought to have been a no. Those marks are the bar earning its keep. The drafts with no mark went out as written or with the wording shuffled about, and each of those stood for a customer kept waiting on a reply that was fine all along.
Left alone, a team counts every edit as a correction, and on that count the bar never comes down, since there is always somebody who would have gone and phrased it their own way. Two months on, the AI is holding back nearly everything and three of you spend your mornings reading drafts. It creeps up on a desk. Somewhere upstairs the person who signed for the AI is asking, not unreasonably, where the time went.
A month of held and sent drafts, read by category
Take a made-up month of 900 tickets at a small online shop, with the AI sending some replies on its own and holding the rest. All the numbers are invented for the example, though they are the kind of numbers you would expect to see. The shop had five kinds of question.
The AI sent 410 replies without anybody touching them. It held 320 as drafts. The last 170 it produced no draft for at all, since it could not make enough sense of what was being asked, and people wrote the whole of it from scratch.
Go through the 320 held drafts with the pen and 214 come back with no mark on them, sent as written or with the wording moved about. The other 106 needed a fact put right, or a proper rewrite. To be honest with you, that is two held drafts in every three that could have gone on their own, which sounds like a bar set far too high, and for some of the shop’s questions it plainly was.
The whole month hides where. Read it by category and it comes apart quickly enough. Order status questions, where is my parcel and has it shipped, the sort that mostly get answered from Shopify, were 320 of the tickets, and the AI held 80 of them, of which 74 were fine as drafted. Somebody went and read 74 correct tracking replies that month, the lot of them a customer waiting. Account and login questions, SSO trouble mostly, went much the same way, 36 fine out of 40 held.
Refunds went the other way altogether. The AI held 50 refund drafts and 30 of them had a fact wrong, usually the amount, which lives in Stripe rather than anywhere the AI was reading, or whether the order was inside the refund window, and of the 30 refund replies it had sent on its own, a person reading 10 of them afterwards in a one-off check found 3 that should never have gone. Nobody went and checked those at the time. On this desk the refund bar wanted raising, or the category taking off the AI’s sending altogether, and the order-status bar wanted bringing down, both in the very same month. Our AI copilot piece puts refunds on the list of things that ought to carry a person’s name every time, next to account closures and GDPR requests, and this month rather bears it out.
Product questions were a different problem again, and we will come back to them further down.
| Category | Tickets | Sent on its own | Held as a draft | Held drafts fine as written | Sent replies read / wrong | No draft at all | What the month says |
|---|---|---|---|---|---|---|---|
| Order status | 320 | 210 | 80 | 74 | 10 / 0 | 30 | Bar too high: bring it down |
| Returns and exchanges | 160 | 70 | 60 | 44 | 10 / 1 | 30 | About right: leave it |
| Account and login | 120 | 60 | 40 | 36 | 10 / 0 | 20 | Bar too high: bring it down |
| Refunds | 100 | 30 | 50 | 20 | 10 / 3 | 20 | Bar too low: raise it, or keep refunds with a person |
| Product questions | 200 | 40 | 90 | 40 | 10 / 2 | 70 | Missing material: write articles before moving the bar |
| All categories | 900 | 410 | 320 | 214 | 50 / 6 | 170 | One number for the whole desk hides all of the above |
How many sent replies should a person read
Nobody has to be reminded to read a held draft, on account of the customer getting nothing until somebody does. The sent replies are another story. Somebody says in week one that they will read a few of those every Friday, and by week three, with a busy Friday or two in between, it has worn off by its self. That is a shame, since the three bad refund replies on the made-up shop were all in the sent pile.
People usually start with a percentage, one in every ten sent replies, say. Fair enough on paper, only the mistakes are not spread out fairly at all. On the made-up shop, one in ten of the order-status replies is 21 of the most boring emails in the business, while one in ten of the refunds is 3, too few to tell you anything. The errors hide in the small categories and a percentage reads them the least.
A fixed number per category works better. Ten a week from each, or every one of them if a category has fewer than ten, read by somebody who knows the policy well enough to spot a wrong date. On the made-up shop that comes to a little under 50 a week, since the refunds would be read in full at 7 or so, and it is an hour or so of somebody’s Friday that reads the refunds as closely as it reads the tracking links. It is not a great deal of reading.
The mark you put on each one is the very same pen mark as before. Would you have changed a fact in it, yes or no. Wording does not count here either, and if you catch your own self marking wording, you are tired, mind you, and it is time to go and hand the reading to somebody else for a week.
There is a second check sitting in the tickets already, and it costs nothing. Somebody writes back on the Thursday about a reply they got on the Tuesday, same ticket, same question, so that reply did not settle it. Go and put the AI’s replies next to your people’s in the one category and count how many of each had a customer writing back inside the week. It takes 10 minutes. If the AI’s count is well up on your people’s, the Friday reading will usually turn up why. The CSAT scores on that category say the same thing in the end, only a month or so after the reopens did.
Why moving the bar only moves the mistakes
Now the product questions, the ones we left for later. The AI held 90 drafts and 50 of them had a fact wrong, it sent 40 and 2 of the 10 read afterwards were wrong, and 70 more tickets got no draft at all. Nudging the bar up on those only fills the held pile with more wrong drafts, and the other way the wrong ones go out instead, no two ways about it, while the reason so many were wrong in the first place has not moved an inch.
Go looking for where the answers to those questions actually live, which on the made-up shop would take the best part of a morning of asking round the office and opening old folders. The size chart turns out to be a supplier’s PDF in somebody’s Google Drive. Whoever does the buying knows the difference between the two waxed jackets, and knows it in their head and nowhere else, which is grand right up until they go on their holidays. The washing instructions went out once, in a Slack message back in March, or it might have been Teams. A couple of the wrong drafts were about a bug the engineers had logged in Jira weeks before, which nobody had thought to mention to the desk. The AI had none of that to draft from, so it was unsure, and it was right to be. It was hardly the AI’s fault.
So on a category like that the lever is not the bar at all. The bigger part of it is writing the answers down, into a knowledge base the AI can draft from, and then watching the no-draft count and the fact-wrong count fall together, which is the only change in this whole business that reduces both kinds of mistake at the one time.
There is a trap in the other direction as well, and our page on the AI copilot goes into it properly, the article that is written down clearly and is simply out of date. Say free delivery went up from 50 to 60 on the first of the month and the delivery article in the knowledge base still says 50. Every draft about delivery will say 50 as well, clearly and confidently, and it will sail over any bar you like, since as far as the AI can tell it has read the rule and applied it. The bar only knows how sure the AI is. It has no way of knowing the article went stale on the first.
Where does the loop sit in Maxdesk
Over here at Maxdesk the AI comes with the Elite plan as two agents, and the one doing all of the above on your desk would be the resolver. A ticket comes in and it writes a full reply from the conversation and whatever sits behind it, and if it is sure enough of that reply, off it goes. Anything under the bar stays in the ticket as a draft for one of your people to read, put right where it needs putting right, and send under their own name, and the full story of how the two agents split the work is over at our AI support agent page.
Whatever either agent does gets written into the ticket’s history, the sending and the holding both, so when you sit down on a Friday with your pen you are reading a record rather than trying to remember a busy week. The usage dashboard puts sent against held category by category, against the month’s allowance of 5,000 AI responses, which is most of that made-up month done for you, apart from the one thing only a person can mark, whether a fact was wrong. That part of it is still yours. We would not trust any tool that claimed to fill it in its self.
Where the resolver keeps holding back week after week on the one sort of question, sizing say, you are looking at an article nobody has written yet, the whole of it sitting in somebody’s head, and going and writing it is the quickest way we know of getting it to hold back less. As for how far you can move the bar on your own workspace, or whether refunds can be kept off sending altogether, put that to us on the day rather than taking it from this page. We would sooner you asked us than planned your month around a setting we have not shown you here.
The best person to read ten refund replies on a Friday is usually whoever handles the refunds in accounts. On our pricing, adding them to the desk for that hour costs nothing, since we charge for the workspace and never count the people in it. We also do not bill the AI by the resolution. A bill that went up every time a reply went out on its own would be quietly leaning on you to drop the bar, and our piece on per-resolution pricing goes into why we will not work that way.
The parts no threshold covers
Every confidence score you will ever see comes from the AI marking its own homework, plainly said. Useful, mind you, and better than nothing by a distance, but it is still an opinion the AI holds of its self. The sampling in this article exists because that opinion needs checking by somebody who knows what the right answer was.
There is no AI at all on our Free or Pro plans, so a desk on either of those has a person behind every reply already, and this page is about a problem it has not got yet. On Elite the AI does all of the above. It works on email, whether the customer writes from Gmail or Outlook, and nothing else, and there is no building of bots of your own, so if what you had in mind was a chat assistant on the website trained up by your own hand, we are not the tool for that.
An allowance of 5,000 AI responses covers a small desk’s month comfortably and a busy desk’s less so, and past that it is packs. Worth watching on the dashboard rather than going and finding out at the month end.
And the reading costs time we have not measured for you, if we are being honest. Fifty or so sent replies on a Friday, plus whatever the bar holds back through the week, is most of somebody’s Friday morning. In the first month, before the bar has been set anywhere sensible, that morning can eat most of what the AI saved. Most of the difference between an agent and a chatbot comes down to who checks what, and we have yet to see one where the answer was nobody.
What should you count in the first month
In the first month we would let the AI draft everything and send nothing, with a person pressing send on every single reply, whatever your tool happens to call that. Not one mistake gets out that way, since every reply has a person on it. When the month is done you have the whole of it by category in a spreadsheet, or a CSV out of your desk if that is easier.
One pen mark a draft, fact changed or not, and nothing fancier, since anything fancier has stopped by the second Friday. That is the lot. By the last day of the month, the categories with hardly a mark in them are the ones you could let send on their own. Where a third or more of the drafts carry a mark, the AI is not ready for that sort of question yet, and on refunds it may never be.
On the made-up shop, a month like that hands over order status and account questions the very next morning, and shows returns to be a closer thing than anybody had assumed. What is left over is a heap of product questions with no written answer anywhere, and that heap is worth an afternoon of your own self before the AI sends a single word.
Common questions about human in the loop and confidence thresholds
What is human in the loop in customer service?
Human in the loop means a person checks an AI’s work at a set point before a customer relies on it. In support that usually means the AI drafts replies and either a person approves every one, or the AI sends the replies it is confident about and holds the rest as drafts for a person.
What is a confidence threshold for AI replies?
It is the level of confidence an AI needs in its own draft before it sends the reply without a person. Replies above the threshold send automatically; replies below it wait as drafts. Raising it means fewer wrong replies sent and more correct ones held back; lowering it does the reverse.
How do you know if the confidence threshold is set too high?
Count the held drafts that went out unchanged or with only wording changes. If most drafts in a category were fine as written, the threshold is holding back correct replies there and costing customers waiting time. In a made-up month of 320 held drafts, 214 needed no factual change.
How many AI-sent replies should a person review?
A fixed number per category rather than a percentage, for example 10 a week from each category, or all of them where a category has fewer than 10. Percentages over-read the common, low-risk questions and under-read small categories such as refunds, where the costly mistakes tend to be.
Should different kinds of ticket have different thresholds?
Yes. The cost of a wrong reply varies by category: a wrong tracking answer costs a follow-up email, a wrong refund answer can cost money. A single threshold for the whole desk will be too cautious for some categories and too loose for others.
Does a high confidence score mean the AI’s answer is correct?
No. Confidence measures how well the reply follows from the material the AI was given. If that material is out of date, the AI can be highly confident and wrong, which is why sent replies still need sampling by a person who knows the current policy.
Reviewed on 29 September 2026. The month of 900 tickets is made up for the example; the Maxdesk AI details were checked on the live site the same day.
