The Hidden Cost of Untested QA in Financial Services


Manual QA costs contact centers millions a year. A bigger risk comes from scorecards that have never been validated against real customer outcomes.
The changing landscape of QA
Call centers resolve customer issues on the first call about 71% of the time, according to SQM Group's industry benchmark, meaning nearly three in 10 customers have to call back about the same problem. The gap compounds from there. When a call isn't resolved on the first try, customer satisfaction drops by roughly 15% for each additional call the customer has to make, and 80% of those customers report being dissatisfied with how the agent handled the call.
Financial institutions are also moving fast on AI. Cornerstone Advisors' 2026 banking survey of senior bank and credit union executives found that 59% of credit unions and 49% of banks have already moved generative AI into production. Across industries, Deloitte Digital's 2026 survey of contact center leaders found that 35% already use agentic AI in operations.
Much of that AI is going straight onto the frontline, on top of a QA process nobody has re-examined in years. The stakes are different at this scale. A person having an off day affects a handful of calls. An AI agent repeating the same mistake can affect the next thousand, instantly, before anyone notices. And in many programs, no one independent is holding the bots to the same standard as the people: they're scored by a different tool, or even graded by the vendor that sells them.
Nearly every contact center has a quality assurance program meant to catch problems before they compound: a scorecard, a sampling process, a calibration routine. By one industry estimate, 92% have a formal QA program in place, but manual review covers only 2% to 5% of conversations for most.
But a more foundational problem goes beyond manual sampling. Most scorecards used to measure quality have never been tested to see whether they predict real customer outcomes. The result is a quality process that looks rigorous on paper, confident on the inside, and, in practice, may be measuring almost nothing that matters to the business.
This report models the cost of manual, low-coverage QA, which runs to millions of dollars a year, and details a potentially bigger risk: a scorecard nobody's ever checked against real outcomes. That second risk is harder to quantify, and it doesn't go away when QA is automated.
Most scorecards used to measure quality have never been tested to see whether they predict real customer outcomes.
Calibrated doesn't mean correct
Coverage is the more visible problem. Reviewing only a sliver of conversations is enough to hide real, costly problems inside the 95% nobody ever looks at.
But fixing coverage doesn't fix the foundational problem. Search for “how to build a call center QA scorecard” and dozens of guides come up, from vendors, associations, and consultancies alike. They largely agree: structure the scorecard into a few sections, weight the criteria, score on a five-point scale, and calibrate until two reviewers scoring the same call reach the same number.
It's good advice. It's also, on its own, solving the wrong problem.
What calibration actually fixes
Calibration closes a real gap: without it, one reviewer's 4 is another's 2, coaching becomes arbitrary, and agents reasonably conclude the whole system is unfair. It's what makes a score mean the same thing no matter who's looking.
But notice what calibration never asks. It never asks whether the criterion itself is worth scoring in the first place. Two reviewers can spend an hour in perfect agreement that an agent scored a 2 on "empathy in tone," and neither one has any idea whether that 2 has ever once correlated with a customer staying, leaving, complaining, or coming back happier.
Most QA programs spend enormous effort making sure a scorecard is scored consistently, and almost none making sure it's scored correctly. Consistency answers “did everyone score this the same way.” Correctness answers a different question: does this criterion, scored perfectly, actually predict anything real about the customer relationship?
Those two questions get conflated, because a well-calibrated scorecard feels rigorous. It has weighted sections, inter-rater reliability, monthly sessions where disagreements get resolved. All of that produces real confidence, confidence that the number is fair, not confidence that the number matters.
Why nobody built the test for this
There are three primary reasons, and none of them are about anyone doing their job poorly.
The first is technical. Testing whether a criterion correlates with a downstream outcome (churn, retention, repeat contact, revenue) means connecting data that typically lives in different systems, owned by different teams: the scorecard in operations or quality, the outcome data in product, data, or finance, and sometimes, all three. Doing that by hand, for every criterion, on an ongoing basis, is hard enough that most organizations never try. They default to what they can measure well: whether the scoring is consistent.
The second is historical. Many of today's criteria were written when QA was primarily a compliance and script-adherence exercise. Did the agent say the required disclosure? Did they follow the mandated steps? That's a legitimate, still-vital function. But as the demands on customer experience evolved well beyond compliance, most scorecards never got redesigned to keep up. They still measure what they were built to measure years ago, not what actually matters now.
The third compounds the other two. Leaders are often measured on the same scores and SLAs the scorecard itself produces. Hitting those targets shows up in their own performance reviews and sometimes on-target earnings. That creates a disincentive to dig deeper: a green dashboard looks good for everyone currently being measured by it, and finding out the criteria underneath it don't predict anything real is not a comfortable discovery to bring to your own leadership.
A similar incentive applies to AI agents. A vendor that defines, reports, and sometimes bills on its own bot's resolution rate has little built-in reason to test whether that number matches what customers experienced.
Nobody has to act in bad faith for this to happen. The incentives just never pointed toward finding out.
A gap that's beginning to close
The industry's thinking is beginning to shift, and the reason is technology, not just awareness. Testing every criterion against real outcomes used to require exactly the kind of manual, cross-system correlation work described above, hard enough that almost nobody attempted it at scale. AI has started to change that, and the conversation is shifting with it. Some newer platforms now caution against building a scorecard around what's easy to measure rather than what actually drives outcomes. Others point out that metrics like CSAT and repeat contact often get added to a scorecard without anyone establishing how they connect to the behaviors underneath them.
What's still missing from most is a full method, a repeatable way to take an existing scorecard, test every criterion against real outcome data at scale, and know which to keep and which to retire. Most of the industry's foundational guidance was written for a world where that kind of testing wasn't practical, so calibration is still often treated as the last step. The technology has moved faster than the playbook.
A different question to ask
None of this means scorecards are the wrong idea or that calibration should stop. Calibration was never the end goal. It was supposed to be step one, with a second, harder step to follow: does this criterion, once everyone agrees how to score it, actually predict the outcome that matters to the business?
That's a different discipline than QA has traditionally practiced, closer to how a product team A/B tests a change before rolling it out, than how a call center builds a rubric: test the criterion against reality, keep what holds up, retire what doesn't, and treat the scorecard as something that gets re-tested, not set once and trusted forever.
The pattern shows up often once anyone actually looks: a criterion everyone assumes matters (tone, phrasing, a specific script line) turns out to move nothing measurable. Meanwhile, behaviors nobody thought to score, such as how an escalation gets handled or what happens in the 60 seconds after a customer says they're frustrated, can matter more than anything the scorecard measures.
AI agents raise the stakes further. A chatbot can be tuned to hit every script line, required disclosure, and polite phrase, so on a typical scorecard it may score near the top. A green dashboard says little about whether the bot understood the question, gave the right answer, or knew when to hand off to a person.
One distinction matters especially in financial services: not every criterion is there to predict outcomes. Required disclosures, identity verification, and other regulatory must-haves stay on the scorecard. The criteria worth testing are everything else, the behavioral and quality measures meant to shape the customer experience.
The stakes of getting this wrong are not small. Financial services firms overestimate how well they handle customer complaints: 60% believe their customers are satisfied, while only 22% of customers actually say they are, according to industry research. A scorecard that's never been tested against outcomes is exactly the kind of system that produces a gap like that and never catches it.
The scorecard isn't the problem. Never checking whether it's right is.
Most QA programs spend enormous effort making sure a scorecard is scored consistently, and almost none making sure it's scored correctly.
The cost of manual QA
The untested scorecard is hard to price, but the cost of running QA manually isn't.
For a typical mid-market contact center, running roughly 250 agents, manual, low-coverage QA costs an estimated $1 million to $1.5 million a year in lost capacity and missed improvement opportunities. For a large enterprise operation, around 2,500 agents, that figure rises to $10 million to $15 million a year, before including major compliance failures, lost sales, or customer churn.
The estimate draws on published staffing benchmarks, compensation data, and a transparent set of assumptions.
per year
250-agent operation
per year
2,500-agent operation
Estimated annual cost of lost capacity and missed improvement opportunities.
How the number is built
The estimate is based on a per-agent figure: lost productivity of approximately $5,000 per contact center agent per year, within a reasonable range of $2,000 to $11,000 depending on interaction complexity, sales exposure, regulation, and the maturity of the existing QA program. At $5,000, a 250-agent operation loses about $1.25 million a year and a 2,500-agent operation about $12.5 million, the midpoints of the ranges.
That estimate is based on a specific set of assumptions:
- One QA analyst for every 30 agents, consistent with published industry staffing ratios of roughly 1:20 to 1:30
- An average loaded QA analyst salary of approximately $79,000: base pay near $59,000 (per salary.com), plus benefits and employer costs
- QA automation capable of removing or redirecting roughly 60% of manual evaluation work, a figure vendors report consistently, though it's worth treating as directional rather than a guaranteed return
- One supervisor for every 15 agents, in line with typical staffing benchmarks, with manual QA administration estimated to consume roughly 12% of supervisor capacity, half of which is recoverable
- Better interaction coverage, faster coaching, and earlier issue detection improving overall agent productivity by approximately 3%
- A loaded annual agent cost of $60,000
These figures cover the cost of running QA manually, but they don't account for whether the criteria being scored are the right ones. Automation alone doesn't fix that.
What testing actually reveals
Nubank, a rapidly growing digital bank serving more than 100 million customers, had built a QA scorecard the way most organizations do: 16 criteria, weighted, calibrated until reviewers consistently agreed on what each score meant. By every conventional measure, it was a mature program.
Until Nubank tested to see if any of the 16 criteria actually predicted anything real about the customer relationship.
It ran each criterion against actual customer outcomes instead of just trusting that consistent scoring meant the scorecard was right. Eight of the 16 held up. The other eight had been scored, coached, and calibrated for years without ever proving anything measurable.
KOHO, a Canadian fintech and digital banking app, found something similar, and reached it a different way. Rather than testing one criterion at a time, KOHO moved to evaluating every single customer interaction and let the pattern reveal itself. It found that grammar mistakes, long assumed to be a serious issue, barely moved the needle at all. What actually mattered was empathy. Coaching shifted accordingly, away from script perfection, toward the human connection that actually drove outcomes.
The result wasn't just a better scorecard. It was a different QA function entirely. Analysts who used to spend their time manually reviewing a small sample now spend it coaching and training, working from what automated full coverage actually surfaces. KOHO's support team has stayed roughly the same size even as its user base has doubled, scale through technology, not headcount.
Put together, the two cases point the same way. Testing the test doesn't just validate scorecards, it redirects them. It surfaces what's actually driving the customer relationship and drops what never moved the needle. Then coaching can follow.
It didn’t take more QA staff or a longer scorecard. It required two things most manual programs never have together: coverage broad enough to see the whole pattern, and the ability to test what that coverage actually predicts.
Testing the test doesn't just validate scorecards, it redirects them. It surfaces what's actually driving the customer relationship and drops what never moved the needle.
Closing the gap
None of this is a failure of effort. Most QA teams are working hard inside systems that were never built to connect scorecards to outcomes, under incentives that reward a green dashboard over an uncomfortable discovery. Nubank and KOHO show that's starting to change.
The playbook those companies followed is repeatable. Inventory the scorecard and set aside the regulatory must-haves. Pick the outcomes the business actually cares about, such as repeat contact, complaints, churn, and retention. Test every remaining criterion against them across full interaction coverage, including outsourced teams and AI agents. Keep what predicts, retire what doesn't, and look for behaviors that matter but were never scored. Then repeat on a schedule, because a scorecard that held up last year may not this year.
In financial services, a scorecard tested against real outcomes is stronger evidence for compliance and risk teams. It shows that quality measures connect to what actually happens to customers, not just that reviewers agreed with each other. It also holds people, outsourced teams, and AI agents to the same outcome standard, scored independently of whoever runs them.
It also matters more with every AI deployment. As AI agents take on more frontline work, the scorecard isn't just grading people anymore. It's grading systems that get tuned to whatever the scorecard rewards, and an untested scorecard can teach an AI agent the wrong lessons a thousand times before anyone notices. The CFPB flagged the risk years ago. It warned that chatbots in consumer finance can fail to recognize disputes and give inaccurate information, and that deploying institutions face compliance risk.
The cost of doing nothing isn't abstract. Manual QA alone can cost an estimated $1 million to $1.5 million a year for a mid-market contact center and $10 million to $15 million for a large financial institution, before accounting for a scorecard nobody's tested against real outcomes.
Financial institutions are sitting on millions of calls a year, all carrying real signals about risk, retention, and what a customer is about to do next. Most of those signals get filed away as a compliance record instead of acted on. When that changes, QA stops being a score one team defends in a review. It becomes something strategic a COO can build resourcing decisions on.
Veja Birdie em ação.
Veja como o Birdie transforma sinais de clientes em decisões de retenção, expansão e adoção. 30 minutos. Demonstração ao vivo com resultados.
