There’s a Reddit thread on r/callcentres that every support leader should read. An agent describes waking up on rollover day, the morning QA scores reset, excited and confident. They’d followed the call flow perfectly on every new booking call they knew would be reviewed.
But QA pulled different calls. The ones where, by their own admission, they don’t bother with the call flow. “Those calls don’t get pulled.” The result: a 52% Internal Quality Score (IQS). “I’m so [expletive] tired of crying over this job.“
What’s striking isn’t the frustration. It’s the strategy embedded in the post. The agent knows which interaction types get reviewed and has adjusted their behavior accordingly. On the calls that count, they perform. On the rest, different standards apply.
That’s not a character flaw. That’s a rational response to a broken system, one where the scorecard has become the goal instead of the tool.
In this article, we’ll cover why QA scorecards quietly become compliance checklists, how that infects reviewers just as much as reps, how to anchor your program to your values instead of your script, and what a practical, outcome-oriented QA scorecard actually looks like.
Why QA Scorecards Become Compliance Checklists
Goodhart’s Law gives us the clearest frame for what goes wrong with QA programs:
“When a measure becomes a target, it ceases to be a good measure.”
As soon as your scorecard includes a “followed script” checkbox, you’ve created a game.
Reps figure this out fast: check, check, check, check, check. Technically they did everything. But something still feels off. The experience isn’t right. The customer’s problem might not even be solved.
Here’s what makes it worse: the checkbox mentality doesn’t just infect reps. It infects reviewers too.
As you scale, with more agents, more volume, more evaluators, the people doing the reviewing face the same pressure. They need to assess X conversations per week.
So they start ticking boxes: “Did they do this? Yes. Did they do that? Yes. Done.” The human judgment leaves the room on both sides of the review.
You’ve dehumanized the process, and as soon as you dehumanize the process, you lose sight of what QA is actually for.
The goal is to drive customer outcomes: customer satisfaction, resolution, the experience of feeling genuinely helped. But when both the rep and the reviewer are focused on checkboxes, that goal quietly disappears.
According to SQM Group’s research, only 17% of agents believe their call center’s QA efforts positively impact customer satisfaction. That’s not an engagement problem. That’s a design problem.
How to Align Your QA Scorecard with Your Team’s Values
The fix isn’t to remove structure from your QA program. It’s to anchor that structure to your values, then translate those values into questions that guide judgment rather than prompts for box-ticking.
When I built the first QA scorecard at the Academy to Innovate HR (AIHR), this was the principle we started from. We were a small team, and our goal was to begin measuring quality in a way that actually reflected what we cared about.
We put our mission at the top: “to set our members up for success for a WOW-quality experience through empathetic communication and purposeful action.” Underneath that, our five values: WOW-Quality, Empathy, Purpose, Integrity, Efficiency. And on the right-hand side of each value, we translated it into a practical question a reviewer could actually use.
“Empathy” became: “Do you use a personal tone that makes our member feel understood, welcomed, and cared for?”
Not: “Did the agent use the empathy phrase?”
The first requires a judgment. The second just requires a box to be ticked.
We also made a deliberate choice about weighting:
- Voice & Tone: 1.2
- Communication Clarity: 1.0
- Solution Accuracy: 1.0
- Follows Aligned Process: 0.8
The category closest to a compliance check sat at the bottom. Process adherence matters, but it isn’t the most important thing an interaction can get right. The weighting communicated that priority to reviewers clearly.
We used a four-point scale with no neutral option. Either the conversation showed the trait, or it didn’t, and to varying degrees. No middle ground. Forcing reviewers off the fence keeps the evaluation honest and prevents lazy scoring.
This scorecard has evolved since those early days, but the founding principle has stayed consistent: connect every scored element to a value, and make sure the question you’re asking the reviewer requires an actual judgment about the customer’s experience, not just a scan for the right words.
When Your Scorecard Rewards Script Adherence Over Judgment
Here’s a concrete example of where most scorecards go wrong.
Say one of your categories is “Voice & Tone.” If the question is “Did the agent follow the voice and tone guidelines?” you’ve got a checkbox. You’ll get agents who sound exactly like the script. You’ll also get agents who are excellent at sounding like the script in reviewed interactions and completely different in everything else.
What you actually want to be asking is something like: “Did the agent stay within our voice and tone while making the experience personal for this specific customer?” That’s a different question entirely.
It requires the reviewer to assess whether the interaction felt human, whether the agent paid attention to the cues the customer gave them, whether the response was right for this situation rather than just for the template.
The same logic applies to any value you want to see in practice. Don’t put “went above and beyond” as a checkbox. Translate it into what it actually means for your team day-to-day.

One of the clearest practical expressions of “above and beyond” in a support context is next issue avoidance: supporting the customer not just with the problem they raised, but with what’s likely to come next.
If someone contacts you to set up a new account, a good agent already anticipates what questions they’ll have next week. They address those proactively. If your scorecard has no way of recognizing or rewarding that, it can’t drive it.
The deeper issue is this: when behaviors are the only thing your scorecard measures, you’re implicitly telling reps that doing the behaviors is sufficient.
As long as they check those boxes, they’re fine. But if you’re actually looking for outcomes, your scorecard needs to be flexible enough to give reps room to make decisions and reward good judgment, not just script adherence.
Why “Solved” and “Resolved” Are Not the Same Thing
There’s a distinction worth making explicit, because it’s where most QA programs quietly go wrong, and it matters more now that AI is taking on more of the support work.
A resolved ticket is one where the process was followed, the boxes were ticked, and the interaction was closed to protocol. A solved ticket is one where the customer’s actual problem is gone and they won’t need to contact you again.
These aren’t always the same. And the gap between them is exactly where human judgment lives.
When a customer contacts support, there’s context surrounding them that goes beyond what they’ve typed or said:
- Frustration they haven’t named.
- Confusion they’re too polite to admit.
- Circumstances where the technically correct answer would land badly because of what’s happening around it.
A human agent reads all of this: voice tone, surrounding circumstances, the cues between the lines. They adapt in the moment.
An AI can read everything in the conversation and deliver the technically correct answer. What it won’t do is reach outside the text to give a contextually right response. That’s a human thing.
Here’s the problem.
If your QA scorecard penalizes agents for adapting their language to a customer who wouldn’t understand the required phrase, or for going slightly off-script because the situation called for it, you’re signaling that you’d prefer the AI’s answer. Clean, compliant, phrase-correct. But not actually solving the customer’s problem in front of them.
This plays out concretely with mandatory language.
If a customer can’t understand the standard phrasing, a good agent will adapt it. Scoring them down for that isn’t a quality outcome. It’s creating the conditions for customer confusion.
Your Customer Satisfaction Score (CSAT) and Customer Effort Score (CES) will pick that up. A compliance-based QA scorecard will say everything was fine.
A QA program that can’t distinguish between resolved and solved ends up optimizing for AI-ready behavior in human agents, and gets neither the consistency of AI nor the contextual judgment of a person.
How to Build Judgment Into Your QA Review Process
One of the most valuable things a QA program can do is create space for agents to explain their decisions.
Not every interaction fits the standard process. Sometimes a customer’s situation is unusual. Sometimes the right answer for this person isn’t the one the script produces.
A good agent will adapt. A scorecard that penalizes them for adapting, simply because they deviated from the checklist, isn’t measuring quality. It’s measuring compliance and calling it quality.
A scorecard oriented around outcomes gives reps room to say: “I made this decision because I believed it was the best outcome for the customer within our guidelines.” That judgment then becomes part of the review, not a mark against the agent for going off-flow.
This is also what separates a QA program that builds capability from one that just enforces compliance. When reviewers are asking “what was this agent trying to do, and did it serve the customer?” that’s a coaching conversation. When they’re asking “did they say the required phrase?” it isn’t.
The hardest part of building this culture is showing people what “good” actually looks like in practice. If nobody has ever seen a conversation that genuinely nails Voice & Tone while staying personal rather than just compliant, they have no reference point to aim for. Calibration sessions should spend as much time sharing positive examples as they do discussing scoring disagreements.
How to Audit Your Current QA Scorecard?
If you’re not sure whether your QA program is measuring quality or compliance, run it through these questions:
- Could an agent score highly by following your scorecard exactly and still leave a customer’s problem unsolved? If yes, your scorecard is a compliance checklist.
- Do the questions on your scorecard require judgment, or just pattern-matching? “Did they use the required greeting?” is pattern-matching. “Did the opening make the customer feel their issue would be handled?” is judgment.
- Are your values visible in your scored elements? If your company values empathy, genuine helpfulness, or ownership, can you point to a specific item that measures each of those things, framed as a question about the customer’s experience?
- Are your IQS scores correlated with Customer Satisfaction Score (CSAT) and First Contact Resolution (FCR) rates? If agents with high quality scores aren’t delivering better customer outcomes, your scorecard is measuring something other than quality. The same diagnostic logic applies to making agent CSAT a useful signal rather than a number to manage around.
- What does a coaching conversation after a low score actually sound like? If it’s mostly “you missed this phrase,” you’re in compliance mode. If it’s “here’s what the customer actually needed, and here’s what could have gone differently,” that’s a quality program.
If your scorecard doesn’t pass this audit, the answer isn’t to add more categories. It’s to go back to your values and rebuild the questions from there.
For teams thinking about setting up a quality program from scratch, Zendesk QA’s foundational course series on YouTube (formerly Klaus, now part of Zendesk) covers the core competencies in a tool-agnostic way. It’s a solid place to start before committing to any specific process or tooling.
One more practical step: your ticket data holds more quality signal than any manual review program captures. Manual QA typically covers just 3–5% of all interactions. Agents figure out the patterns quickly and optimize for them.
Swifteq’s Zendesk ChatGPT app can automatically extract sentiment and outcome signals from ticket content across your full volume, surfacing patterns your manual sample is almost certain to miss.
Build a QA Program That Measures Outcomes, Not Compliance
Goodhart’s Law will always apply. Make a number the target, and people will find a way to hit it, even if the path there undermines everything the number was supposed to represent.
The agents posting in support communities about hating their QA scores aren’t disengaged. They’re often the most perceptive people on the team. They’ve read the system and responded rationally. When your best performers are gaming the scorecard, the scorecard needs to change.
A QA program that measures outcomes looks different from one that enforces compliance. It asks harder questions, produces better coaching conversations, and over time builds a team that genuinely solves problems rather than one that’s learned to look like they do.
If you’re using Zendesk and want to get more from your quality data, start your free 14-day trial of Swifteq today, no credit card required. Or schedule a demo and we’ll show you what’s possible.
What changes when the sample is every ticket
The agent in that Reddit thread wasn’t lazy. They were rational. They knew which interactions got pulled and performed on those, and the rest ran on different standards. That strategy only works because manual QA sees three to five percent of the volume. The other ninety-five percent, where most of the real quality signal lives, is where nobody is looking.
Close that gap and the game changes. When every closed ticket is evaluated, there is no reviewed subset left to optimize for. The calls that count and the calls that don’t become the same calls, because they all count.
That is what ResolveLoop does. It reads every closed Zendesk ticket and asks the question this article is built around: was the customer’s problem actually solved, or just closed to protocol? And when the answer is no, it names why. Sometimes the agent went off-script for the right reasons and the script itself was the problem. Sometimes the agent did everything well and the real cause was the product, the policy, or the process around the conversation. A compliance check can’t tell those apart. Outcome-oriented evaluation across the full volume can.
Grading three percent of tickets for script adherence measures compliance. Reading all of them and asking what actually happened to the customer measures quality.
Written by Neal Travis
Curious learner and builder of customer experiences in scale-ups. Neal is the Head of Customer Experience at the Academy to Innovate HR (AIHR) and Host of Growth Support.




