Quality

Support QA Scorecard: Six Criteria, Three Points, One Automatic Zero

9 Sep 2026·16 min read

A support QA scorecard is a fixed set of criteria you score a sample of finished conversations against, so quality gets measured on what you sent rather than on how the week felt. The one below has six criteria, scores each of them 0, 1 or 2, weights them to 100, carries one automatic zero, and asks a second question about every failure that most scorecards never ask.

That second question is the reason this page exists, so we may as well put it up front. When a reply scores badly, somebody has to decide whether the agent could have done better with what they had in front of them at the time. A good half of the time they could not, and scoring it against them anyway is how a QA programme quietly dies around week six with nobody admitting they stopped.

What is a support QA scorecard, and what is it for?

It is a rubric, with a sampling habit to go along beside it, and that is genuinely the whole of it.

You take a handful of conversations that are already finished with, you read them properly rather than skimming down them, you score each one against the very same criteria every week without changing your mind halfway, and then you write down what you found. The output is not a league table and it was never meant to be one. What comes out is a short list of things to go and teach, and a shorter list of things to go and fix in the desk its self, and the second of those lists is usually the more valuable one by some distance, though it rarely feels that way at the time your own self.

What the whole of it is for is the gap between what your customers tell you and what your team actually sent them. CSAT asks the customer how it felt, which is well worth having, mind you, and is also weighted heavily toward the outcome rather than the craft, on account of a customer who got their refund tending to rate the reply kindly whether or not it was any good at all. Your QA score asks somebody on the inside whether the reply was correct and complete and clear. The two of them disagree constantly, and all the interesting findings live inside of that disagreement.

Why do most support scorecards stop being used by week six?

Because of a failure mode nobody writes about, and it is worth naming before you copy anything off this page.

The usual shape of the thing runs like this. Somebody goes and builds a rubric with fourteen criteria on it, each of them scored out of 5, covering tone and empathy and grammar and product accuracy and formatting and follow-up and heaven knows what else besides. Week one everybody is keen as anything and the reviews take an hour. Week three the reviewer is skimming. Nobody says so. By week five the scores have all settled into the very same narrow band and not one person can remember what a 3 means as against a 4, nor could they explain it to a new starter if you asked them to. And by week six the review has slid to the bottom of somebody’s Friday list, on account of it producing nothing anybody acts on, and there it sits until a new manager arrives and builds a fresh rubric with fifteen criteria on it.

Three things did that between them. The rubric was too long for anybody to hold in their head. The scale was too wide to use honestly. And the findings had nowhere to go except into a conversation with the agent, which meant every finding felt like a complaint about a person, and people go and stop generating findings that feel like complaints. Fair is fair, that is a rational response to the whole business rather than laziness on anybody’s part.

So the scorecard below is short on purpose, narrow on purpose, and routed in two directions on purpose.

The scorecard, filled in

Six criteria, each one scored 0 or 1 or 2, and weights adding up to 100. Copy the whole of it into a sheet. Change the weights if your business genuinely needs different ones, though we would keep accuracy sitting heaviest whatever else you move about.

#Criterion2 points1 point0 pointsWeight
1AccuracyEverything said was true and completeTrue but incomplete, customer would need to ask againAnything factually wrong. AUTOMATIC ZERO for the whole conversation30
2Answered the actual questionAddressed what was asked, including the unasked part behind itAddressed the literal question onlyAnswered a different question, or sent a link instead of an answer25
3Clear enough to act onCustomer knows exactly what happens next and by whenClear, but no next step or no timeframeCustomer has to interpret it, or there are three possible readings15
4Tone fits the situationMatched the customer's state, warm without being performativeNeutral, template-ish, nothing wrong with itDefensive, dismissive, or cheerful at somebody who is upset15
5Handled properly behind the scenesRight tags, notes a colleague could pick up cold, correct statusDone, but thin notes or wrong tagsNo notes, wrong status, colleague would start from nothing10
6Closed the loopConfirmed the fix landed, or told the customer when it wouldClosed with a reasonable assumption it workedClosed with the thing still open, or left hanging5

And then there is the column that actually matters, and it sits beside the whole conversation rather than beside each criterion.

Could the agent have done better with what they had in front of them? Yes or no. If a reply lost points because there was no article to send, or no permission to issue the refund, or no template, or no way at all of finding out the answer inside the time available, then the answer is no and the finding belongs over on a backlog rather than in a coaching note. Anybody can make that judgement for them selves after a fortnight of practice. It is no subtle call. The honest answer is usually plain enough inside about 30 seconds of reading, truth be told, and the reviewer knows it before they have finished the thread.

Weighted scoreWhat it meansWhat to do about it
90 to 100The reply you would want on your own accountNothing. Keep one of these a month as a teaching example
75 to 89Solid. Something small was missingNote it, mention it in the next one to one, no action beyond that
60 to 74The customer got by, but it cost them effortCoach, or fix the desk gap, depending on that second column
Under 60Rework territory. Reach back out to the customer if it is recentCoach and fix, and check whether the same gap is hitting others
Any accuracy zeroWrong information went outCorrect it with the customer today, then work out how the wrong answer was available to be given

Why score in threes rather than out of five?

Because five-point scales collapse into fours, and one and all who have run one know it well enough, and very few of them ever write it down.

Give a reviewer 1 to 5 and go and watch what happens to it over a month or two. The 1s vanish first, on account of nobody much wanting to write a 1 about a colleague who is going to read it on Monday. Then the 5s go the very same way, seeing as a 5 feels like saying nothing at all could be improved, and there is always something. So what you are left holding is a scale of 3 and 4 with the odd rare 2 in it, which is a two-point scale carrying three extra numbers about with it for decoration, and it produces averages that look precise on a slide and move up and down for no reason anybody in the room can explain.

Three points forces the call. Either the thing was done, or it was half done, or it was not done, and each of those is a judgement a person can make quickly and defend afterwards. It also makes calibration between two reviewers a great deal easier, since disagreeing about a 1 against a 2 is a conversation you can finish, while disagreeing about a 3 against a 4 is a conversation about feelings.

The nuance goes back in through the weighting instead. Accuracy at 30 and closing the loop at 5 says plainly enough what your desk cares about, and that ratio is a better statement of your standards than any amount of scale precision was ever going to be.

What should be an automatic zero?

One thing only, and that is accuracy.

If the reply contained something untrue, the conversation scores zero, no matter how warm it was and no matter how quickly the thing went out the door. That is harsh and it is meant to be harsh. A wrong answer delivered beautifully does the customer more damage than a blunt right one, seeing as they will go and act on it, and the acting on it carries the whole of the cost, in a refund that never comes or a setting they change on a Tuesday that breaks something else by Thursday.

Resist adding any more automatic zeros than the one. Every one you add is a criterion you have quietly decided matters as much as accuracy does, and once there are four of them sitting there the scorecard stops discriminating altogether, and every bad conversation ends up looking the same as every other bad conversation. Tone should not be an automatic zero, mind you. Missing tags certainly should not. Those cost the business something, no two ways about it, and neither of them costs a customer a decision made on bad information.

The other side of it, and this is the part that keeps the rule survivable at all, is that an accuracy zero has to be treated as a system question rather than as a hanging offence. Somebody gave a wrong answer, so go and look for the place that wrong answer was sitting, available to be given, in the first place. An old article nobody ever retired. A colleague who told them the wrong thing in perfectly good faith. A product change that shipped on a Wednesday without one person telling support about it. Go and find that before anybody has a difficult conversation with the person who typed the reply.

How do you tell an agent problem from a desk problem?

This is the column nobody else has got, and it is far and away the most useful thing on this page, so it is worth spelling out how the judgement actually gets made in practice.

Read the conversation through, then go and ask what the agent had in front of them at the moment they replied. Was the answer written down anywhere at all in your knowledge base? Did they hold the permission needed to do the thing the customer was asking about, or would they have had to go off and find somebody who had it? Was there a template for this, and if there was one, did it fit the situation or was it written for a different sort of customer entirely? Could they see the account history, the order sitting in Shopify, the payment in Stripe, the CRM record, whatever the question happened to need, without leaving the thread and asking about it in Slack?

If the answer to that lot is broadly yes and the reply was still weak, then it is a coaching finding and it goes to the one to one. If the answer is no, then the thing you have found is a desk problem with an agent’s name attached to it by accident, and the fix for it is an article, or a permission, or a template, or a change to the way something gets routed at tier 1, or now and again a bug that needs raising in Jira and chasing. Write it on a backlog and move along. Nobody needs a conversation about that one.

Now, what to do with the pair of numbers, and this is the bit worth setting up properly from week one rather than bolting on in March. Count both columns every week and keep an eye on the ratio between the two of them. A desk running mostly coaching findings has a training gap on its hands. A desk running mostly backlog findings has a tooling and documentation gap instead, and coaching people harder will change nothing whatever about it, though it will go and make one and all tired. Most teams running this for the first time are fairly surprised by how heavily the thing tilts toward the backlog side, and that surprise is the whole return on the exercise, truth be told.

There is a quieter benefit to the two columns as well, which is what the whole arrangement does to the room. Agents will accept being scored a good deal more readily once they can see that half the failures get routed away from them, and a reviewer who has to justify which column a finding lands in ends up reading the conversation far more carefully than one who is only assigning a number to it.

How many conversations should you actually review?

Fewer than you would think, and here we part company with most writing on the subject, which quietly implies you are measuring something.

For a team of four, review 5 conversations per agent per week, which is 20 of them in total. That is somewhere around an hour of somebody’s time at roughly 3 minutes a conversation once the reader has the hang of it, and nearer 2 hours in the first fortnight while everybody is still arguing about what a 1 means. Pull them at random rather than picking out the interesting ones, on account of the interesting ones being known about already, and the whole of the value here sitting in the ordinary conversations nobody would ever have gone and looked at. Export the week’s finished conversations to a CSV and pick with a random number if you want it done honestly, and give up the pretence that scrolling and picking is random, because it never was.

Now the honest part. Twenty conversations out of a week where the desk handled eight hundred of them is a sample of 2.5%, and 2.5% will not tell you your quality has moved by 4 points, whatever the average at the bottom of the sheet says. What it tells you is what a handful of replies looked like, which is genuinely useful for teaching and genuinely useless as a measurement. Treat the weekly number as a coaching signal and nothing grander. Treat any month-on-month movement under about 10 points as noise, seeing as that is what the whole of it is.

Two more sampling rules worth having about the place. Review conversations that are already closed, so you are scoring the whole of an exchange rather than half of an unfinished one. And spread the picks across the week rather than taking the lot from Monday, seeing as a desk on a Thursday afternoon is a different desk altogether to the same desk on a Monday morning, and the difference between the two of them shows up in the replies.

Does QA tell you anything CSAT and the other metrics do not?

Yes, and the disagreements between them are the useful part of it rather than the awkward part.

CSAT is the customer’s own verdict and it gets shaped by the outcome more than anything else. Somebody who got what they wanted rates the interaction well even when the reply was a mess of a thing, and somebody who was told a firm and perfectly correct no rates it poorly even when the reply was as good as your desk has ever sent. Your scorecard is blind to the outcome and reads the craft instead. Put the pair of them side by side over a quarter and a pattern turns up that is worth acting on. Conversations scoring well inside and badly with the customer are usually policy problems rather than support problems. And the reverse of it, a high CSAT sitting on a low-scoring reply, is usually a customer who was easy to please, along with a habit that will cost you dearly the first time it meets a harder one.

The speed metrics sit differently again. First response time and average handle time, FRT and AHT on most dashboards, measure the clock and know nothing whatever about the content, and both of them can be improved handily by sending something quick and useless. QA is the only one of the whole family that reads what was actually said, which is the reason a desk with a spotless SLA record and no QA at all can go on shipping fast rubbish for months on end without a single number moving on anybody’s report.

None of that makes QA the important one. It makes it the missing one on most desks, which is a different claim and a smaller one.

How do you run it, week by week?

The whole of it wants to cost about 90 minutes a week, and it wants to be a boring 90 minutes at that, on account of the boring version being the one still running in March.

  1. Friday morning, pull the week's closed conversations and pick 5 per agent at random, using a sheet rather than your judgement. Not the escalated ones, not the ones somebody complained about, and not the ones you happen to remember.
  2. Score each against the six criteria, 0 or 1 or 2, and let the sheet do the weighting. Do not stop to write essays. A short note on anything scoring 1 or 0 is plenty.
  3. On every finding, mark the second column. Could the agent have done better with what they had, yes or no. That judgement takes seconds and it decides where the finding goes next.
  4. Split the output into two lists. Coaching findings go to the one to one. Backlog findings go on the list of articles to write, permissions to widen, templates to build and routing to change.
  5. Once a month, have two people score the same 3 conversations separately and then compare, which is calibration and takes about 20 minutes. Where the two of them disagree, it is the criterion wording that needs fixing rather than the reviewer.
  6. Once a quarter, look at the ratio of coaching findings to backlog findings and at nothing else. That single ratio tells you whether your quality problem is people or plumbing, and it is the number to bring to anybody holding a budget.

Then go and leave the rubric alone for a quarter. Rubrics that get edited every month produce numbers nobody can compare against last month’s, and comparing against last month was most of what you wanted the numbers for in the first place.

Where our own product stands on this, and what it will not do

Plainly said, because a page that hands out an artifact is a poor place to start selling.

Maxdesk has no QA module. There is no scoring field, no review workflow, no calibration tool and no quality dashboard anywhere in the product, and we are not going to describe a roadmap we have not built. If you run this scorecard you will run it in a spreadsheet beside the desk, the same as nearly everybody else does.

What the product does have is the material the review needs, which is not nothing. Every conversation in the shared inbox carries its full history, the internal notes your team wrote along the way, the tags, the statuses and the audit trail of who touched what and when, so a reviewer reads the whole thing through rather than piecing it back together out of Gmail or Outlook, wherever your mail sits, Google Workspace or Microsoft 365, along with somebody’s memory of a Slack thread from a fortnight ago. There is a read-only role, which is the honest way to give a reviewer or a team lead access without handing them the ability to change anything. And you can tag reviewed conversations, so finding last month’s sample again is a filter rather than an archaeology project. Those are workarounds, mind you, and we would rather call them that than dress them up.

On the money, since it changes what a QA habit costs to run. We charge per workspace and never per seat, so adding a reviewer, or a team lead who only reads, adds nothing to the bill. That is the difference that matters here: on per-seat pricing, the person doing your quality reviews is a line item, and a QA programme that costs money per reviewer is a QA programme somebody will cancel. Free is $0, Pro is $20, Elite is $99, all with every agent included, and the history you can see runs 3 months, 12 months and 24 months in that order. The free plan also carries ads and a small branding line on outbound mail. Checked 7 September 2026.

The case against scoring your team at all

We have handed you a rubric, so it is only right to hand you the arguments against using it.

The strongest one is Goodhart, and it applies here properly rather than as a slogan. Tie these scores to pay or to a ranking and the scores will improve while the support gets worse, since agents will write to the rubric, which rewards a complete and clear and well-tagged reply and cannot tell whether the customer’s actual problem went away. We would go a good deal further than most and say that if you cannot commit to keeping the whole of this out of compensation, then do not start it at all. A QA programme that has gone wrong does more harm than never having run one, seeing as what it teaches your team is that somebody reading their work is a thing done to them rather than for them, and that lesson takes a long while to unlearn.

Next after that, and we said it further up and will say it again here because it is the thing most likely to get forgotten the moment somebody puts the weekly average on a slide, a small sample is not a measurement. Twenty conversations will tell you what those twenty conversations looked like and they will not tell you a great deal about the other seven hundred and eighty.

Then there is our own bias, which is worth stating plainly on our own site. A page arguing that half your quality failures are missing articles and missing permissions is a page arguing for a desk with a knowledge base and permissions in it, which is a thing we sell. Read the second column argument with that in mind, and then go and run it for a fortnight, because the ratio it produces is yours and not ours and it will settle the question either way.

And the plain limits. This scorecard was built for written support. We ship email and no other channel, so a desk running phone or live chat will want different criteria for the parts that happen live, and the escalation path in particular scores differently when it happens on a call. The weights on the table are a starting point rather than a finding, since we have no traceable research saying accuracy ought to be 30 rather than 35. And there are no benchmark scores anywhere on this page on purpose. Nobody publishes a QA average that means anything across different rubrics, and any figure claiming otherwise was made up somewhere along the line. We re-check this page in the first quarter of 2027, weights and bands and all.

Common questions about support QA scorecards

What should be on a customer support QA scorecard?
Six criteria at most: accuracy, whether the actual question was answered, clarity, tone, how the ticket was handled behind the scenes, and whether the loop was closed. Weight accuracy heaviest, at around 30 of 100, and make anything factually wrong an automatic zero for the whole conversation.

How many tickets should you review per agent?
5 per agent per week works for a small team and costs about 90 minutes in total. Pull them at random from closed conversations rather than choosing the memorable ones. It is a coaching sample, not a measurement, and treating a 2.5% sample as a metric is the commonest mistake made with it.

Should QA scores affect pay or performance reviews?
No. Tie a rubric to money and people write to the rubric, which improves the scores and not the support. Use the scores for coaching and for finding gaps in your documentation and permissions, and if you cannot keep them out of compensation, it is better not to run the programme at all.

What is the difference between QA scores and CSAT?
CSAT is the customer’s verdict and it follows the outcome, so a refund earns a good rating whatever the reply was like. A QA score is an internal read of the craft and ignores the outcome. They disagree often, and the disagreements usually point at a policy problem rather than a support problem.

How do you stop QA feeling like surveillance?
Route half the findings away from people. Mark every failure with whether the agent could have done better using what they had, and send the ones they could not to a backlog of articles, permissions and templates. Agents accept review far more readily once they can see that gaps in the desk get counted against the desk.

Do you need QA software to run a scorecard?
No. A spreadsheet with six columns and a weighting formula runs this perfectly well, and Maxdesk has no QA module to sell you anyway. What matters more than the tooling is a reviewer with read access to full conversation history and a habit that survives past week six.