Guide

Chatter Performance Scorecard: The Grid, Ready to Copy

A chatter performance scorecard you can copy: seven criteria, a shared scale, and the volume-band correction that stops raw revenue ranking the wrong person.

Published 28 July 2026

7criteria on the scorecardstructure of this guide
1-5scale used on every criterionstructure of this guide
roughly a thirdof conversion lost when the pitch lands before the sixth messageour data
38,879sales behind these findingsour data

Every agency eventually builds a leaderboard, and every leaderboard eventually ranks the person who was handed the best inbox. The scorecard below exists to stop that. It scores the work, not the fans someone inherited.

Copy the grid, keep the scale, and resist the urge to add up the column.

What is a chatter performance scorecard for, and who uses it?

It turns “how is she doing” into seven answerable questions, asked the same way every month. That is the whole purpose: comparability across people and across periods.

  • The owner uses it to allocate shifts. The peak evening should go to the highest timing and closer scores, not the highest revenue line.
  • The team lead uses it as a coaching agenda. One low criterion is one conversation to have, with examples attached.
  • The chatter uses it to know what good looks like before the review, not during it.

It is not a ranking, and it is not a pay formula. The moment it becomes either, people start writing for the grid instead of for the fan.

What does the complete scorecard look like?

Seven criteria, one scale, no total. Score each from 1 to 5, where 1 is “costs us threads”, 3 is “safe unsupervised”, and 5 is “I would hand them a difficult account”.

# Criterion What you are reading Score 1 Score 5 Source
1 Conversion at comparable volume Conversion rate inside their volume band Bottom of the band Top of the band Dashboard + band split
2 Pitch timing Where in the thread the pitch lands Pitches in the first few messages Waits, and the wait is deliberate Read conversations
3 Closer quality The last line of a sales message Ellipsis, or a closed question at the close Open, specific, no trailing dots Read conversations
4 Voice discipline Drift from the creator’s reference messages Own vocabulary surfacing Indistinguishable from the reference Read conversations
5 Floor-price discipline Behaviour under a haggle Discounts, or invents a bundle Refuses and keeps the thread alive Read conversations
6 First-response time on live threads Median wait on threads inside their block Threads outlive the block Answered inside the block, every time Dashboard
7 Handover hygiene The note left at the end of the block Missing, or unusable Next shift never re-pitches a purchase Handover notes

Recorded but not scored, in a separate row of the same sheet:

Period: ____________          Chatter: ____________
Threads handled: ______       Volume band (1 / 2 / 3): ______
Blocks worked: ______         Creators covered: ____________
Gross sales attributed: ______   ← recorded, never scored
Conversations sampled for this review: ______

And the review note, written before the meeting, not during it:

Scorecard | [chatter] | [period] | reviewer [____]
Scores: 1)__ 2)__ 3)__ 4)__ 5)__ 6)__ 7)__
Strongest criterion, with one example thread: ____________________________
Weakest criterion, with one example thread: ____________________________
One thing to change before the next period: ____________________________
One thing to keep doing: ____________________________
Compared with last period, what moved: ____________________________

How do you score each criterion?

Four of the seven need conversations read end to end. Pull a sample across different blocks, including one quiet afternoon and one peak evening, so you are not scoring the easy hours.

  • Criterion 1. Split the roster into thirds by threads handled, then read conversion only inside a third. Never across bands.
  • Criterion 2. Count the messages before the pitch. In our corpus, pitching before the sixth message drops conversion by roughly a third, and the optimum sits after about ten exchanges (the reasoning is in when to pitch a sale).
  • Criterion 3. Look only at the final line. Two habits cost you conversion: the ellipsis is the worst-performing closer we have measured, and a closed question at the moment of closing costs several points of conversion. See how to end a sales message.
  • Criterion 4. Read five consecutive replies against the creator’s reference set. Drift shows up in punctuation and jokes before it shows up in vocabulary.
  • Criterion 5. Find one haggle. There is always one.
  • Criterion 6. Median, not average. One catastrophic thread should not decide the score.
  • Criterion 7. Read the note the next shift received, not the one that was written.

One more habit belongs in the coaching conversation even though it does not carry its own row: a personal callback placed at the exact moment of the pitch costs points, even though it feels like it should help. It is the most counter-intuitive result in our corpus and people do it instinctively: the personal callback mistake.

What do you record about revenue instead of scoring it?

Three lines, all of them context and none of them scored: gross sales attributed, threads handled, and the volume band those threads put the person in. Revenue belongs in the sheet (a review that cannot see the size of the period is not a review), but it belongs above the grid rather than inside it, because it measures the inbox and the grid is trying to measure the person. The longer argument for that is in creator agency metrics that matter; the scorecard’s job is narrower, which is to keep the number in the room without letting it rank anyone.

What you compare What it actually measures Ranks correctly?
Gross sales Who inherited the best fans No
Sales per hour worked Inbox quality, again, per hour No
Conversion, whole roster Volume differences between people No
Conversion inside a volume band How well comparable work was done Yes
Messages before the pitch A habit our data ties to conversion Yes, as a leading signal

The correction is one step: build the bands before you build the ranking. Comparing someone working cold threads with someone holding long-standing whales produces a number that looks precise and means nothing.

What are the three classic mistakes?

Averaging the seven into one score. A chatter scoring 5 on closer quality and 1 on floor discipline averages to something unremarkable, and the average hides the only fact worth acting on. Read the profile, never the mean.

Changing a criterion mid-period. Comparability is the entire value of the grid. If a criterion is wrong, note it, finish the period, and change it at the quarterly revision.

Scoring from the dashboard alone. Criteria 2, 3, 4 and 5 are invisible in aggregate numbers. An agency that scores only what its tooling counts ends up rewarding volume and speed, which are the two things that push people to pitch early.

When do you revise the scorecard?

Revise the grid quarterly, and run the review monthly. Weekly reviews measure noise; quarterly reviews arrive too late to change anything.

Three triggers justify an off-cycle revision:

  1. Two reviewers cannot agree. Rewrite the criterion until two people reading the same five conversations land within one point of each other.
  2. A criterion stops separating anyone. If everyone scores 4 or 5 on it, it has become a house standard. Move it to the onboarding checklist and free the row.
  3. The work changes. A new platform, a new creator with a different audience, or a change in what the team is asked to sell all move what “good” means on rows 2 and 5.

Keep the old grids. The first thing anyone asks in a review is what moved since last time, and that answer only exists if the criteria stayed still.

Frequently asked questions

Why not just rank chatters by revenue?

Because the scorecard scores work, and revenue is not work. It is context. It goes in the recorded block at the top of the sheet, next to threads handled and volume band, so a reviewer can see the size of the period without a number ranking anyone. What gets scored instead is criterion 1: conversion inside the volume band, where the inbox is held roughly constant between the people being compared.

What is a volume band and how do I set mine?

Split your roster into thirds by threads handled in the period, and compare conversion only inside a third. You are not looking for an industry benchmark, you are looking for the spread between people doing comparable work. The bands come from your own roster, so they move as the team changes.

How long does scoring one chatter take?

Budget the time to read a sample of their conversations end to end, because five of the seven criteria cannot be read from a dashboard. Sample across different blocks, including one weekday afternoon and one peak evening, so you are not scoring the easiest hours.

Should the scorecard drive pay?

Keep them apart. The moment a criterion sets pay, people optimise the criterion rather than the conversation. The easiest one to game is the one that costs you most, the early pitch. Use the scorecard for coaching and shift allocation; keep pay on terms agreed in the contract.

What if two people score the same chatter differently?

That is the grid failing, not the reviewers. Rewrite the criterion until two people reading the same five conversations land within one point. A criterion nobody can score consistently is a criterion that produces arguments instead of coaching.

Does this work for a solo creator with no team?

Yes, scored against yourself over time rather than against anyone else. Run the same seven criteria on your own threads once a month. The timing and closer criteria are the ones that move fastest, and they move for a creator exactly as they do for a chatter.

See what it looks like in practice

The justonedash chatbot holds the conversations, keeps each creator’s voice and works around the clock.