quality control
Quality control in chatting is the regular review of a random sample of real conversations, scored against written criteria, to correct what chatters say to fans before revenue shows the damage.
Quality control is not reading dashboards. A conversion rate tells you a shift is selling less than it did; it never tells you why, and it takes weeks to say even that much. Quality control opens the conversations themselves and hands you the cause on the first read, off a tiny sample. It is a habit with a fixed slot in the week, not an investigation launched when a month goes bad.
How often should you read conversations back?
Weekly, at a fixed time, around ten conversations per chatter. Rhythm beats volume: a monthly review arrives after the habit has already set in.
Two sampling rules stop you from fooling yourself.
- Pull at random, not from threads that sold. Winning threads only show you what already worked.
- Take at least one thread from every shift, nights included. A team reviewed on office hours only never gets read on the hours that carry the month.
Be honest about the ceiling. Ten conversations a week per person is a sliver of what actually goes out. Quality control samples; it does not monitor.
What do you actually look at?
Checkable things, not impressions. Split the sheet in two and only score the left column.
| Countable | Arguable |
|---|---|
| Which message number carried the offer | “the thread felt cold” |
| How the sales message ended | “he could have pushed harder” |
| Price stated once, in plain figures | “that wasn’t quite her voice” |
| Thread picked up correctly at clock-in | “I’d have done it differently” |
The first two rows carry most of the value. Our data puts an offer before the sixth message at roughly a third of conversion lost, with the optimum after about ten exchanges. The ending of a sales message matters on its own: a closed question at the moment of closing costs several points of conversion, and an ellipsis is the worst-performing closer we have measured. All three are visible at a glance, with no argument about taste.
How do you score without demotivating?
Score the move, never the person, and separate what comes from the chatter from what comes from the shift.
- A written grid, known in advance. Nobody discovers the criteria on scoring day.
- One mark, one screenshot. A criticism with no excerpt attached is not a criticism, it is a mood.
- Compare like shift with like shift. A 4am thread is not judged against a 9pm one: conversion collapses between 2am and 6am and peaks in the evening whoever is typing.
Skip the third and the score reads as unfair, which is the point at which people stop reading it.
Who should run it?
Someone whose paid hours include it, not the agency owner whenever a gap opens up. Quality control that depends on management having a free afternoon disappears in the first busy month, exactly like the handover.
What comes out of a review does not get fixed thread by thread. It goes back upstream into the voice guide and into how the shift rotation is cut, which is what stops the same mistake happening twice.
Related terms
To go further on this:
Read the full guide