Most Zendesk quality assurance (QA) programs run the same way: a team lead pulls a handful of closed tickets each week and scores them against a scorecard, usually somewhere between 2% and 5% of total ticket volume.

The other 95%-plus closes without anyone looking at it.

Your QA program has an opinion about a tiny, non-random slice of your volume, and nothing to say about the rest. When something does get flagged, the default assumption is almost always that the agent got it wrong, even when the real cause never touched the agent at all.

In this article, we’ll cover how Zendesk QA got here, what a good customer service QA scorecard needs to get right, how auto QA changes the coverage math, and the practical steps to actually set it up, including how to tell whether a low score is an agent problem or a structural one before you act on it.

How Zendesk QA Evolved: from Manual Review to Auto QA

For years, dedicated QA software was a separate purchase from your helpdesk.

Teams bought it specifically to add structure, and eventually AI scoring, to their review process.

Klaus was the best-known example for Zendesk teams: an independent, AI-powered QA platform that Zendesk acquired in early 2024 and rebranded as Zendesk QA.

That mattered less as an acquisition headline and more as a signal: full-coverage, AI-scored QA stopped being a specialist add-on and became a native part of the platform most Zendesk customers already run their support on.

In my experience at the Academy to Innovate HR (AIHR), our own QA process moved through a similar arc: a manually reviewed Word document, then a standardized spreadsheet, then an AI-assisted first pass before a human made the final call.

Each step cut down how long it took to get a read on quality. None of them, on their own, changed how much of our ticket volume anyone actually saw. Auto QA is the step that does.

What Does a Good Customer Service QA Scorecard Actually Measure?

A customer service QA scorecard is an evaluation framework that breaks a support conversation down into specific, weighted categories, so a reviewer (human or AI) can score it the same way every time instead of going on gut feel.

The trap most scorecards fall into: process categories like Spelling and Grammar or Greeting carry the same weight as whether the agent actually solved the customer’s problem.

That tells your team a well-punctuated non-answer scores about the same as a genuine fix.

There’s more to scorecard design than this section can cover: what to weight, how to phrase categories so they require judgment instead of pattern-matching, and how to avoid rewarding script adherence over outcomes.

We’ve written a full breakdown of how to build a scorecard around outcomes instead of compliance, and a related piece on avoiding scores that misattribute blame to the agent.

Both are worth reading before you finalize your categories. For the rest of this guide, the one thing to take from scorecard design: whatever your categories are, at least one needs to ask whether the customer’s actual problem got solved, not just whether the agent followed the right steps.

How Does Auto QA Changes the Coverage Math?

Even a well-designed scorecard is only useful if it’s applied to more than a token sample.

The 2-5% ceiling from the intro doesn’t move as your ticket volume grows.

A 40-agent team handling a few thousand tickets a week produces far more conversations than any reviewer can read, so the sample shrinks as a percentage precisely when you need it most.

Industry benchmarks put manual QA coverage in that same range regardless of team size, and the tickets a random sample is least likely to catch are usually the ones revealing something structural, like a policy generating the same complaint across dozens of agents.

Auto QA closes that gap by scoring the ticket instead of sampling around it.

In Zendesk QA, the AutoQA feature scores every closed or solved conversation automatically, human and AI agent alike, against whichever categories you’ve activated, plus any AI-prompt categories you build on top.

Coverage goes from a percentage to effectively all of it, though it’s worth knowing AutoQA only evaluates new conversations going forward, not your historical backlog, and needs at least ten words from both customer and agent to score a ticket at all.

The more useful shift is what a full score makes possible.

Zendesk QA’s Spotlight feature, for example, uses the full ticket volume to automatically flag the conversations most worth a human’s attention: churn risk, stalled threads, knowledge gaps, escalations.

A reviewer stops guessing which 3% to read and gets pointed straight at where the actual risk is sitting.

Getting Started with Zendesk QA: a Practical Setup Checklist

If you’re moving from a manual sample to full coverage, here’s a working order of operations rather than trying to flip every switch at once.

Turn on AutoQA with at least one outcome-focused category active.

Beyond the defaults (Spelling and Grammar, Greeting, Closing, Empathy, Tone, Readability, Comprehension, Solution), you can build up to 20 custom exact-match categories and up to 10 AI-prompt categories, where you describe what to look for in plain language.

Make sure Solution (or your own version of “did we actually fix it”) is one of the categories you activate, not an afterthought.

Turn on Spotlight before you try to read everything yourself.

Spotlight flags the conversations most worth attention: churn risk, escalations, knowledge gaps, stalled threads. It covers all of your coverage, not a sample, so a reviewer’s limited time goes to the tickets that actually need it.

Tag every ticket by intent and topic as it comes in.

Ticket Classification groups your AutoQA scores by ticket type automatically, which is what turns “these ten tickets scored low” into “refund tickets are scoring low across nine different agents.

That’s the signal you actually need to catch a structural issue instead of coaching nine people individually.

Check the pattern before you schedule the coaching conversation.

One agent with a low score is worth a 1:1. The same low score across a dozen agents on the same kind of ticket is usually one of three things: a macro nobody’s written yet, a policy forcing an answer customers don’t want, or a product bug nobody’s fixed.

Fix the system before you coach the person.

Decide what still needs a human read.

AutoQA only scores conversations going forward, not your historical backlog, and needs real substance from both sides to score a ticket at all. Short tickets, historical audits, and anything ambiguous still belong to a person.

    If you’re not ready for full AutoQA coverage but want to test the idea first, Swifteq’s MCP (Model Context Protocol) Server for Zendesk is a lighter starting point.

    Point an AI model at a batch of closed tickets and check them against your own criteria: which ones show the agent skipping troubleshooting steps, which ones were marked resolved but still show negative sentiment.

    It covers far more volume than a manual sample ever could, well before you’re ready to commit to always-on scoring.

    Run through this once, and you’ve gone from a 3% guess to full coverage with a built-in way to tell agent issues from system issues. Those are two things a manual sample could never give you at the same time.

    Building a Zendesk QA Program That Scores the System, Not Just the Agent

    Zendesk QA has come a long way from a shared Word document and a handful of manually reviewed tickets.

    If your program is still running on that model, the honest problem isn’t effort, it’s math. Sampling 3% of your volume was never going to catch the pattern sitting in the other 97%.

    Once AutoQA, Spotlight, and ticket-level tagging are actually running, QA stops being a monthly report on a handful of agents and starts pointing you toward what’s worth acting on:

    • training programs where a theme repeats across enough tickets to justify one,
    • individual coaching where a pattern really is isolated to one person,
    • customer or product issues that nothing else in your stack tracks as consistently as your ticket data already does.

    Get the setup right, and your QA program becomes what it should have been all along: a live read on where your support system is actually breaking, and where to fix it next.

    For more on building a smarter approach to Zendesk QA, start your free 14-day trial of Swifteq today, no credit card required. Or schedule a demo and we’ll show you how Ticket Classification and our MCP Server can fit into your QA workflow.


    Neil

    Written by Neal Travis

     

    Curious learner and builder of customer experiences in scale-ups. Neal is the Head of Customer Experience at the Academy to Innovate HR (AIHR) and Host of Growth Support.


    Similar Articles

    Hiring for the Queue AI Is Building for You

    Hiring for the Queue AI Is Building for You

    There's a version of the AI-in-Support conversation that goes like this: AI handles the simple stuff, humans handle the complex stuff, everyone wins. It's a reasonable starting point. It's also incomplete enough to produce some genuinely bad hiring decisions. The...