A customer service QA scorecard is a rubric that scores a sample of real customer conversations against weighted, written criteria. A typical scorecard has 10 to 20 line items grouped into three or four categories, each weighted by importance, plus a short list of auto-fail items that zero the whole evaluation. The output is a percentage score per conversation and, over time, a coaching map per agent.
Support teams measure speed obsessively and quality almost by accident. Average handle time, first response time and backlog are all in the dashboard by default, because the ticketing system counts them for free. Conversation quality is not counted for free. Someone has to read the conversation and make a judgment, and unless that judgment is structured it is just an opinion with a number attached.
That is what the scorecard is for. It turns "that reply was a bit blunt" into a specific, repeatable, defensible evaluation that a manager and an agent can both look at without arguing about taste.
What is a QA scorecard?
A QA scorecard is a standardized evaluation form used to review individual customer conversations. Each line item states one observable behavior, such as whether the agent confirmed the customer's identity or set an expectation for follow-up. A reviewer scores each item, weights are applied, and the conversation gets a single score. Scorecards exist so that two different reviewers grading the same call land in roughly the same place.
The word "scorecard" gets used for two different things in support, and mixing them up causes real confusion in meetings. A QA scorecard grades one conversation. A team or department scorecard tracks aggregate KPIs like CSAT, resolution time and backlog over a month. This article is about the first one. The second is covered in the rundown of customer service metrics and KPIs worth tracking.
What should a customer service QA scorecard measure?
Nearly every workable scorecard sorts its criteria into the same three or four buckets. The naming varies by vendor and consultant, but the content does not.
| Category | What it asks | Example line items |
|---|---|---|
| Resolution and accuracy | Did the customer's problem actually get solved, correctly? | Correct diagnosis, accurate information given, complete answer, no unnecessary transfer, follow-up committed and completed |
| Communication | Was it clear, correct and appropriately human? | Grammar and spelling, plain language instead of jargon, acknowledgment of the customer's situation, tone matched to the context, active listening or careful reading |
| Process and compliance | Did the agent follow the rules the business is bound by? | Identity verification, required disclosures, correct disposition and tagging, data handling, refund or credit authority respected |
| Efficiency and handling | Was the customer's time respected? | Reasonable hold and response gaps, no repeated questions the customer already answered, correct routing on first attempt |
The line items in the first category are the ones worth arguing about, because that is where scorecards most often go wrong. A scorecard that gives 40 points to greeting, empathy statement and closing script, and 10 points to whether the customer's issue was resolved, will train your team to be charming and useless. Weight the outcome heavily, or the rubric will quietly optimize for theater.
Write criteria as observable behavior, not adjectives
"Showed empathy" is not scoreable. Three reviewers will score it three ways and every disputed evaluation will turn into a debate about feelings. "Acknowledged the impact on the customer before moving to troubleshooting" is scoreable, because you can point at the sentence or its absence. Every line item on a good scorecard passes the same test: could two reviewers who dislike each other still agree on whether it happened?
A customer service QA scorecard template
Here is a working starting point for a general support team. It is deliberately short. Long scorecards do not get filled in.
| Criterion | Category | Weight | Scale |
|---|---|---|---|
| Identified the customer's actual issue, not just the stated one | Resolution | 15 | 0 to 5 |
| Provided accurate, complete information | Resolution | 15 | 0 to 5 |
| Resolved on this contact, or set a specific next step with an owner and a date | Resolution | 15 | 0 to 5 |
| Used clear language the customer could act on | Communication | 10 | 0 to 5 |
| Acknowledged the customer's situation before problem solving | Communication | 10 | 0 to 5 |
| Tone appropriate to the customer's state (frustrated, confused, routine) | Communication | 10 | 0 to 5 |
| Spelling, grammar and formatting fit for the channel | Communication | 5 | Yes or no |
| Followed verification and disclosure requirements | Compliance | Auto-fail | Pass or fail |
| Tagged and dispositioned the contact correctly | Compliance | 5 | Yes or no |
| Stayed within refund, credit and concession authority | Compliance | Auto-fail | Pass or fail |
| Did not make the customer repeat information already provided | Efficiency | 10 | 0 to 5 |
| Routed or escalated correctly on the first attempt | Efficiency | 5 | Yes or no |
That comes to 100 weighted points plus two auto-fail gates. Adapt the wording to your product, but resist adding line items until you have removed one. Every criterion you add costs reviewer minutes on every single evaluation, forever.
How should you weight QA scorecard criteria?
Weight by consequence. Ask what a failure on each line actually costs the business, then let the weights follow. Resolution and accuracy items usually deserve 40 to 50 percent of the total because a wrong answer creates a second contact, a refund, or a churned customer. Communication typically lands at 25 to 35 percent. Process and compliance items are often better handled as auto-fails than as weighted points.
Industry changes the shape. A regulated contact center in banking, insurance or healthcare will push compliance up hard, sometimes to 35 or 40 percent of the weighted score on top of the auto-fail gates, because a missed disclosure is a fine rather than an inconvenience. A product support team for a technical tool will push accuracy up instead. A retention or billing team will weight the concession and authority items more heavily than either.
One useful discipline: after you set the weights, take five conversations you already have strong opinions about and score them. If your best conversation does not score highest, your weights are wrong, not your instinct.
Scoring scales: binary, Likert and auto-fail
Most mature programs use all three rather than picking one.
- Binary (yes or no). Best for items that either happened or did not: the ticket was tagged, the disclosure was read, the follow-up email was sent. Fast to score, almost impossible to argue with.
- Graded scale (0 to 3 or 0 to 5). Necessary for anything with degrees, like tone or clarity. A 0 to 5 scale needs written anchors for at least the top, middle and bottom, or reviewers will drift toward 3 out of politeness.
- Auto-fail. Reserved for non-negotiables: skipping identity verification, giving out data to an unverified caller, exceeding refund authority, being abusive. An auto-fail sets the whole evaluation to zero regardless of the rest. Keep the list short. If you have nine auto-fails, none of them mean anything.
Add a "not applicable" option to every line item and make sure your scoring math excludes N/A items from the denominator rather than counting them as zero. This sounds trivial. It is the single most common arithmetic bug in homemade scorecards, and it silently punishes agents who handled short, simple contacts.
Call center quality assurance scorecard vs chat and email
The categories carry across channels; the line items do not. A call center quality assurance scorecard scores things that only exist on a call: hold procedure, dead air, verbal verification, the closing. Chat and email scorecards score written mechanics, response gaps within a live session, and whether the agent kept the customer informed while researching.
| Channel | Criteria only that channel needs | What drops off |
|---|---|---|
| Voice | Hold and dead-air handling, verbal verification, call closing and recap | Formatting, link accuracy, canned response fit |
| Live chat | Response gaps inside the session, concurrent-chat quality, use of saved replies without breaking context | Hold procedure, verbal tone |
| Email and tickets | Complete answer in one reply, subject line and formatting, correct next-step commitment | Real-time pacing, dead air |
Run one scorecard per channel and keep the category weights aligned so the scores stay comparable. If your saved replies are doing a lot of the work, they belong in the QA conversation too; a weak template will drag down scores across every agent who uses it, which is a content problem rather than a coaching one. Our library of customer service response templates and the guidance on writing customer service scripts that do not sound scripted are the place to fix that.
How many conversations should you review?
Small enough to sustain, large enough to be fair. Most teams land between two and six evaluations per agent per month. Below two, one bad day defines an agent's quarter. Above six, reviewer time becomes the bottleneck and the program quietly dies within two quarters.
Sample deliberately rather than randomly. A blend works better than pure random selection: a portion random for baseline fairness, a portion targeted at contacts that already look interesting (reopened tickets, low CSAT responses, long handle times, escalations, refunds above a threshold). The targeted sample is where you learn things. The random sample is what makes the scores defensible.
Review reopened tickets in particular. A reopen is the cheapest available signal that the first contact did not actually resolve anything, which is the same failure that shows up in your first contact resolution rate.
Calibration: keeping scores consistent between reviewers
An uncalibrated scorecard produces numbers that mean nothing. Two reviewers scoring the same conversation 92 and 68 is common in the first months of a program, and agents notice immediately.
The fix is a standing calibration session, usually monthly, sometimes fortnightly while a program is new. Everyone who scores conversations reviews the same two or three contacts independently, then the group compares scores line by line and argues out the gaps. The point is not to agree on a number. The point is to find the criteria whose wording allows two honest people to reach different conclusions, and rewrite them.
Track reviewer variance the same way you track agent scores. If one reviewer is consistently eight points harsher than the rest, that is a scorecard and training issue, and it is your problem before it is the agents'.
QA score vs CSAT: why they disagree
They measure different things and they should not match perfectly. A QA score measures whether the agent did the job correctly according to your standards. CSAT measures how the customer felt about the outcome, which includes things the agent does not control: the price, the policy, the wait before they reached anyone, and the product defect that started the whole thing.
A conversation can score 100 on QA and 1 on CSAT because the agent flawlessly delivered a "no" the customer hated. The reverse happens too, when a warm, likeable agent gives a confidently wrong answer that the customer will discover next week. Both patterns are useful. Persistent high QA with low CSAT usually means your policy or product is the problem, not your team. Persistent high CSAT with low QA usually means your scorecard is measuring the wrong things. The mechanics of the customer-side measure are covered in the guide to CSAT surveys and how to calculate the score.
Common QA scorecard mistakes
- Scoring compliance with a script instead of quality of outcome. If your top-weighted items are greeting, brand phrase and closing, you are auditing recitation.
- Too many line items. Twenty-five criteria means each is worth four points, which means nothing on the card is clearly important.
- Using the scorecard as a disciplinary instrument. The moment scores drive pay or performance plans without a coaching layer in between, agents optimize for the rubric and stop taking useful risks with customers.
- Never revising it. A scorecard written for a product you shipped two years ago is grading behavior that no longer matters. Revisit it twice a year.
- No feedback loop back to the agent. A score delivered without a conversation is a grade, not coaching, and it changes nothing.
Turning scores into coaching
The scorecard is an input to a conversation, not the conversation itself. The useful pattern is narrow and specific: pick one criterion where the agent scores consistently low, show two real examples, agree on one behavior to change, and re-review in two weeks against the same criterion. One thing at a time beats a fourteen-line report every month.
Aggregate scores are worth reading horizontally as well as per agent. When the whole team drops on the same criterion in the same month, the cause is almost never the team. It is a policy change, a new product, a broken macro, or a knowledge base article that went stale, which is why keeping the knowledge base accurate and current does more for average QA scores than any individual coaching session. And when the pattern points at process maturity rather than individual skill, scoring the operation itself with a structured organizational assessment is a more honest diagnostic than another round of agent reviews.
Frequently asked questions about QA scorecards
What is a good QA score?
Most contact centers set a passing threshold between 85 and 90 percent, with anything below 80 triggering a coaching plan. The absolute number matters far less than its stability. A team averaging 78 on a strict scorecard is in better shape than one averaging 96 on a scorecard that only checks whether the greeting was used.
Who should do QA evaluations?
A dedicated QA analyst produces the most consistent scores, because scoring is a skill that improves with volume. Team leads scoring their own agents is workable at small scale but drifts, since leads grade the people they coach. Peer review adds useful perspective and is a strong development exercise, but it needs tighter calibration than either of the others.
Should QA scores affect agent pay?
Link them cautiously if at all. Tying scores directly to compensation reliably produces gaming: agents steer toward short, easy contacts and reviewers soften scores to avoid conflict. Most teams get better results using QA for coaching and development, and reserving formal performance consequences for sustained patterns rather than individual evaluations.
Can AI do QA scoring?
Increasingly, for part of it. Automated systems can now check every conversation rather than a two percent sample, and they are reliable on objective items: whether a required phrase appeared, whether the ticket was tagged, response gaps, sentiment trend. They are weaker on judgment items like whether the agent identified the customer's real problem. The sensible arrangement is machine scoring for coverage on the objective criteria, human review on the judgment ones and on anything the machine flags.
How is a QA scorecard different from a customer service SLA?
An SLA is a promise about response and resolution timing that you make to the customer. A QA scorecard is an internal standard about how the work is done. You can hit every SLA target and still handle contacts badly, which is exactly why teams that only track SLA compliance are surprised when satisfaction slips. The two belong side by side, and the timing side is laid out in the breakdown of customer service SLA examples and metrics.
Start with ten criteria and one calibration session
The programs that survive are the ones that start small enough to actually run every week. Ten line items, three categories, two auto-fails, two evaluations per agent per month, and a calibration session before anyone sees a score. Add sophistication once the habit exists. A modest scorecard that gets filled in beats an excellent one that lives in a spreadsheet nobody opens.