There’s a version of the AI-in-Support conversation that goes like this: AI handles the simple stuff, humans handle the complex stuff, everyone wins. It’s a reasonable starting point. It’s also incomplete enough to produce some genuinely bad hiring decisions.
The framing itself isn’t a problem. It’s that “complex” is doing too much work. When teams try to operationalize it, they end up with something like: complex means multi-step, or technically difficult, or requiring product expertise.
That leads to agents getting trained heavily on product knowledge, rewarded for fast and accurate resolution, and the assumption that the relational side of things will sort itself out through personality fit.
If that’s your team, you might now be wondering why the AI-deflected queue is fine and your human queue keeps generating the escalations you were trying to prevent.
What’s actually left for human agents once AI handles the tractable volume isn’t just “harder” tickets. It’s a qualitatively different category of interaction. Understanding what makes that category different changes what you should be hiring for.
The Conversations AI Sends You
When AI handles an interaction well, it’s because the interaction had a specific property: the customer was in a cooperative orientation.
They had a bounded question, a straightforward request, a specific problem they wanted resolved. They were oriented toward information exchange. The conversation was recoverable toward resolution from the first message.
Researchers who study conversation use the term tractability to describe this quality, though the basic idea is intuitive to anyone who’s worked a Support queue.
Some conversations are recoverable from the opening line. Others are already in difficulty before you’ve said a word.
AI handles the tractable ones well.
What it routes to your human agents, by design and by failure mode, is everything else: the conversations with low entry tractability.
The customer who’s contacting Support for the third time. The one who’s been with you for four years and is “just wondering about their options”. The one whose message fragments into three different complaints at once. Or the one who’s been burned by a commitment someone in Sales made that your Support team can’t honor.
These customers aren’t arriving in a cooperative orientation. They’re arriving in a repair-seeking one. And that changes the entire job.
The First Quality: Relational Repair Before Information Exchange
In a tractable interaction, the job is to exchange information correctly. Get the facts right, be clear, give the customer what they need to act.
There’s a framework I developed called the Conversational Integrity Model, and it names these as CRIT failures when they go wrong: Clarity, Relevance, Informativeness, Truthfulness.

AI can be trained to avoid a lot of CRIT failures. It’s getting better at this.
But there’s a structural prerequisite that CRIT requires, and it’s one AI cannot supply – at least not at the required rate.
Before information exchange can work, both parties have to be in a cooperative orientation. If the customer isn’t there yet, technically correct information lands wrong.
Some examples of how this shows up in your queue:
- Customer asks for a refund. The bot replies with a statement on the refund policy. A technically correct answer, but the customer hears your organization refusing to acknowledge their experience.
- Customer asks for the feature they were promised. The AI explains the timeline clearly, and the customer feels chased in circles.
The responses given in these examples are relevant to the literal messages and completely miss what the customers were actually trying to achieve.
What has to happen first, in any interaction with a customer who isn’t cooperatively oriented, is relational repair.
The agent needs to demonstrate genuine recognition of what went wrong and why it matters to this specific person. Not a formula. Not “I’m sorry to hear you’re having trouble”.
Something that proves, through its specificity, that the agent has actually processed the customer’s particular situation.
This is what Erving Goffman called face-work: the social labor of restoring someone’s sense of dignity after an event that damaged it.
Everyone who’s worked Support knows what good face-work feels like from the practitioner side. You read the message, understand what’s underneath it, and write something that meets the customer where they actually are rather than where the ticket category says they are.
It’s not a script. It can’t be a script, because what makes it work is precisely the evidence that the specific person was seen.
The quality you’re hiring for here isn’t empathy in the generic sense.
It’s something more specific: the capacity to read an emotional state accurately, calibrate a response to match it, and demonstrate recognition without being either performative or clinical.
Cognitive scientists would call this emotional granularity: the ability to make fine distinctions between emotional states rather than collapsing them into broad categories like “frustrated” or “upset”.
An agent who can distinguish between a customer who is angry and one who is defeated, and who knows those require different responses, is exercising a skill. It’s trainable. It’s also assessable, which is important if you’re trying to hire for it rather than hoping it shows up by accident.
The Second Quality: Reading What Customers Are Doing, Not Just What They’re Saying
The most consistent failure pattern in automated Support isn’t that AI gets the facts wrong. It’s that AI answers the question asked rather than the question meant.
A customer who writes “I’ve been with you for three years and I’m just wondering what my options are at this point” is not asking for a list of account options.
They’re issuing a retention signal, in the politest possible register, giving the organization one last opportunity to respond before they leave. An automated response that produces the account options list has committed a relevance failure so complete it’s almost architectural. Technically responsive. Completely wrong.
In linguistics, the gap between what a message says and what it’s doing is called illocutionary force. “I suppose I could just cancel” is not a statement of intention. It’s a negotiating move. “This is the third time I’ve had to contact you about this” is not a factual report. It’s an escalation demand embedded in neutral language. “I don’t really understand what happened” is not a comprehension gap. It’s a trust repair request.
Human agents who are good at their jobs read illocutionary force accurately.
They don’t respond to the surface of the message. They respond to what the customer is trying to do with it.
And in an AI-deflected queue, the interactions that make it through to humans disproportionately involve indirect communication, embedded signals, and the kind of language customers use when they’re frustrated but still trying to be polite.
The quality you’re hiring for here is sometimes called active listening, which is accurate but undersells what’s actually happening.
A better description might be illocutionary accuracy: the ability to read what a message is doing, not just what it’s saying. Agents who do this well often describe it as a felt sense rather than a cognitive process. They can’t always explain how they know the customer is about to cancel. They just know.
The good news is that this skill, while hard to pronounce, is not hard to assess. Pull a set of real customer messages and ask candidates what the customer is actually trying to accomplish. The gap between agents who see it and agents who don’t is visible in about thirty seconds.
The Third Quality: Judgment When Rules Don’t Cover It
The third category of interaction that AI cannot handle well is the one where the right answer in this specific situation is different from the standard answer.
Rules and policies are meant to address the general, overall anticipated use cases. The customer in front of you has a set of circumstances that the policy wasn’t necessarily designed for.
An agent who applies policy correctly in these situations has done their job correctly in a narrow sense, and failed it in a broader one.
The customer whose billing error falls outside the standard refund window, who has been patient for two weeks through a complicated process, who has a four-year history with the company, probably deserves an exception.
Not because exceptions are always right, but because the capacity to recognize when they are is itself a professional competency.
In the Conversational Integrity Model, this sits under what I call Discretionary Judgment: the ability to identify the right action in a situation that no policy fully anticipated, and to act on that identification without waiting for authorization.
It’s derived from the philosophical concept of practical wisdom:
The stable disposition to orient skilled action toward what’s actually good rather than what’s procedurally compliant.
You cannot script this. You can give automation rules, and good automation should have a lot of them, but you cannot give it the recognition of when rules don’t apply.
That recognition requires something rules can’t encode: the capacity to receive a specific situation as it actually is, rather than as a token of a category the policy was written to cover.
What you’re hiring for here is sometimes visible in how candidates talk about past decisions. Not “I followed the policy” or “I escalated to a manager“, but evidence of having seen a situation, formed a judgment about what it actually required, and acted on that judgment.
Agents who have done this, and who can articulate what they noticed and why it mattered, are the ones who will handle the post-AI queue well. Agents who are waiting to be told the right answer in each situation will struggle, because the right answer in the interactions AI can’t handle isn’t in the playbook.
What This Means for Hiring
The competencies above aren’t new. Every experienced Support practitioner reading this has been exercising them for years. What’s new is that they’re no longer distributed across a full ticket queue. They’re concentrated in it.
When AI handles your cooperative, tractable volume, you’re left with a human queue that runs hotter, carries more relational complexity, and requires more consistent judgment than any queue you’ve managed before.
The agents who handle that queue well are the ones who can do relational repair before information exchange, read illocutionary force accurately, and exercise discretionary judgment in situations policy didn’t anticipate.
The hiring question that follows from this isn’t “does this person have good communication skills”. It’s three more specific questions:
Can they demonstrate they’ve read the emotional state of a message accurately and calibrated a response to match it, in writing, under realistic conditions?
Can they tell you what a customer was actually trying to do in an indirect or embedded message, and what that required from them in response?
Are they able to describe a situation where the standard answer wasn’t the right answer, articulate what they noticed that made it different, and explain what they did instead?
Those three questions assess for the qualities that matter in the queue AI is building for you. The candidates who answer them well are worth the cost of the package you’re about to offer them. The queue they’ll be managing earns it.
If your team uses Zendesk, Swifteq can help automate some of your customer support queue, making your customer service team more effective across the board.
A Different Kind of QA for That Queue
The hotter the queue runs, the less a conventional QA score explains. Once AI clears the tractable volume, what’s left is the queue this article describes: repair-seeking customers, embedded signals, situations no policy anticipated. On those tickets an agent can read the emotion right, answer what the customer actually meant, and still land a bad outcome, because the policy was wrong for the case, or because Sales promised something Support could never honor.
Conventional QA scores the agent against a rubric. It can’t tell you whether a bad outcome came from the agent’s handling, the product, the policy, or a process upstream of the queue. So it keeps coaching agents on problems they didn’t create.
ResolveLoop is a different kind of QA. It reads every closed ticket, diagnoses what happened, and attributes the cause to one of four sources: the agent, the product, the policy, or the process. You find out whether the queue is running hot because of who you hired, or because of decisions made above the queue.
Built for Zendesk teams today. See how it works at resolveloop.app.
Schedule a demo today to learn more!

Written by Ines van Dijk
Ines van Dijk is the founder of Customer Support Excellence and the author of The Customer Support QA Playbook. Her research on the Conversational Integrity Model is indexed on SSRN.



