Quality

From sampling to every call, with humans still in calibration.

Scoring every conversation changes what quality teams can see. It only works if the scorecard stays yours.

September 22, 2026 · 5 min read

For most of the history of the contact center, quality assurance has meant sampling. A QA analyst listens to a handful of calls per agent per month, scores them against a form and shares the results in a coaching session. It is careful work, done by skilled people. It is also a very small window onto what actually happens on the floor.

What sampling misses

A small random sample is good at estimating an average. It is poor at finding the things that matter most. A missed disclosure on one call in a few hundred is a real compliance exposure, but a sample may never catch it. A new agent who struggles only with one call type may look fine on the calls that happen to be selected. A process problem that frustrates customers every afternoon can hide behind a steady monthly score.

Sampling also creates a fairness problem. Agents know their score depends on which few calls were picked. A single bad call in a small sample can dominate a review, while an agent with a quiet pattern of shortcuts may never be seen.

What changes when every call is scored

Automated QA reads every conversation against your scorecard. On Bill Gosling Voice, Auto QA scores 100% of conversations, both the AI agent calls and the human ones. That shift changes three things.

  • Compliance risk is found the same day. A missed mini-Miranda or a disclosure read late is flagged on the call where it happened, while there is still time to act.
  • Coaching starts from patterns, not anecdotes. Supervisors can see that an agent consistently skips confirming terms back, rather than debating one recording.
  • Operations sees drivers. When every call is scored, you can see which call types, times or policies are generating repeat contacts and effort.
spine·voice // auto qa · calibration setcall #61145_
Auto QA · Collections scorecard v8sample
84machine
81QA lead
1disagreement
v8scorecard
CriterionResultWeight
Mini-Miranda read at open✓ pass
Empathy shown on hardship mentiondisputed
Confirmed terms back to customer✓ pass
calibration note added · rubric v9 draftedReview
Calibration review, sample data · rendered by spine·voice

The risk: a machine that grades itself

Full coverage brings a new risk. If the automated scorer drifts, misreads a criterion or applies a rule too literally, the error is now applied to every call rather than a few. Agents and supervisors quickly lose trust in a score they think is wrong, and once trust is gone, coaching stops working.

The answer is not to go back to sampling. It is to keep people in charge of what the machine is measuring, and to check its judgment regularly. That is what calibration is for.

How to keep humans in calibration

Calibration has always been part of good QA: analysts score the same calls independently and then compare, so everyone applies the form the same way. With automated scoring, the machine joins that process as one more scorer. The practices we use look like this:

  • Your scorecard, versioned. Criteria, weights and auto-fail rules come from your QA team. Every change creates a new version, so you always know which rubric scored which call.
  • A regular calibration set. QA leads score a set of calls on a regular cadence, and the platform compares their scores with the machine’s criterion by criterion.
  • Disputes go to a person. Any agent or supervisor can dispute a score. The dispute routes to a QA lead, whose decision stands and is recorded.
  • Disagreements improve the rubric. Where the machine and people disagree consistently, the fix is usually a clearer criterion. The rubric is updated, tested and released as a new version.
  • Golden calls guard against drift. A fixed set of reference calls with agreed scores is rerun whenever the scorer or the model changes, so a regression is caught before it reaches the floor.
The machine joins calibration as one more scorer. It does not get the final word.

What QA teams do with the time

Moving to full coverage does not remove the need for QA analysts. It changes their work. Less time goes into listening to random calls and filling in forms. More goes into reviewing flagged conversations, running calibration, handling disputes, refining the scorecard and working with supervisors on coaching plans. Those are the parts of the job that need experienced judgment, and they are the parts that improve performance.

Scores also feed directly into coaching. A supervisor can open the exact moment a call turned, see the criterion that was missed and assign a practice scenario with a simulated customer.

Agents benefit too. When a score comes from every call rather than a handful, a single bad day no longer defines a review. The conversation between agent and supervisor moves from arguing about one recording to agreeing on a pattern and a plan to fix it.

Starting the move

A practical path is to run automated scoring alongside your existing sample for a period, compare results openly with your QA team and agents and only then make the automated score the primary measure. Start with objective criteria like disclosures and verification, where agreement is easiest, before moving to judgment criteria like empathy.

The bottom line

Scoring every call gives quality and compliance teams a complete picture for the first time. Calibration keeps that picture honest. Keep the scorecard in human hands, give people the final word on disputes and treat the scorer like any other version-controlled system, and full coverage becomes something your floor trusts.

← All articles
Book a demo

Score your own calls against your own scorecard.

Bring a scorecard and a few sample recordings. We will score them live and walk through a calibration review with you.