How to Test What Works in Fan Conversations: Protocol
A test protocol for agencies: one variable, comparable time slots, honest assignment, and the chatter bias that quietly decides most results before you start.
Published 28 July 2026
Most agencies have run something they called a test this year, and learned nothing they can rely on. The idea was usually fine. The setup was not: two versions handled by different people, on different fans, in different hours, stopped on the day one of them looked good.
This page is a protocol, not a result. It is how you find out whether a change to your conversations actually works, on your own inbox, without fooling yourself. Our own corpus of 42,852 conversations could never be run this way: nobody randomised anything, so what we publish from it is an association, which is precisely why it is a hypothesis for your inbox and not a setting to copy.
Why does most agency testing produce nothing you can trust?
Because the two groups were never comparable in the first place. The analysis is not where tests break; the assignment is.
Four confounders decide most agency tests before a single message goes out:
- The chatter. The senior chatter who has been on the account longest is holding the deepest threads and the biggest spenders. Give them the new version and it wins.
- The fans. Threads that are already warm convert on almost anything. A version tested on established relationships is not being tested at all.
- The hours. Evenings and dead nights are not interchangeable. Our data has a measured finding on that alone, covered in what time fans buy.
- The moment. Payday weeks, a viral post, a creator’s promo run. Any of them will move your number more than the sentence you rewrote.
If the two arms of your test differ on any of those, the result belongs to the confounder, not to your change.
What do you write down before the test goes live?
Six lines, on one sheet, before anything is sent. Filling this in afterwards is not the same exercise. It lets you pick the reading that suits you.
| Field | What goes in it | What happens if you skip it |
|---|---|---|
| The variable | The one thing that differs, written as two exact alternatives | You change three things and cannot attribute the result |
| The population | Which threads are eligible: stage, fan type, creator, platform | Warm threads leak into one arm and win it |
| The assignment | How a thread lands in A or B, decided by rule, not by preference | Chatters route the fans they like to the version they like |
| The window | Start date, end date, in whole weeks | You stop on the day the numbers look right |
| The outcome | The single number you decide on, named in advance | You shop around for a metric that agrees |
| The decision rule | What result makes you adopt, revert, or extend | The test ends in a discussion instead of a decision |
Print it. A test that does not fit on that sheet is not ready to run.
How do you change only one variable when a conversation has dozens?
By testing a position, not a whole style. Define the smallest unit that can be swapped without dragging anything else with it.
- Testable: the closing line of a sales message. The order of teaser and price. Whether the offer names a number or a range. The wording of the first follow-up.
- Not testable as one variable: “a warmer tone”, “better scripts”, “the new training”. Each of those is a bundle, and a bundle result tells you nothing about its parts.
- Hold everything else fixed. Same price, same media, same stage, same voice guide. If the price moved too, the test is dead.
The chatting scripts guide covers what is worth putting on the list. This page is only about proving which entry on that list is right.
How do you stop your best chatter from deciding the result?
Assign at the level of the thread, never at the level of the person. That single move removes the most common bias in this trade.
Three ways to do it, in descending order of quality:
- Alternate threads within each chatter. Every chatter runs both versions, so skill, shift and inherited fans sit on both sides. This is the default.
- Split by an arbitrary property of the fan. Odd and even account IDs, for example: anything the chatter does not choose and cannot see.
- Rotate versions across whole shifts. Weakest, acceptable only when the tooling gives you nothing better, and only if you rotate enough times that no one shift pattern owns a version.
What never works: giving version B to whoever volunteers, or to your strongest chatter “so it gets a fair trial”. The volunteer effect and the inherited-whale effect both push the same way, and both are larger than most wording differences.
How do you keep the hours comparable?
Run both versions at the same time. Sequential testing (old version this month, new version next month) compares your change against the calendar.
- Same clock, same days. Both versions live across the same evenings, the same weekends, the same dead hours.
- Whole weeks only. A window that starts mid-week has compared different parts of it.
- Check the split afterwards. Count how many threads in each arm fell in evening hours. If one arm got more of them, your randomisation leaked and the result is about your rota, not about your wording.
How much volume is enough, and when do you call it?
Enough that removing your single largest sale from the winning arm does not change who won. That is a test you can run in a spreadsheet, and it is worth more than a significance calculation nobody in the room can interpret.
- Do not peek. Look once, at the end of the window. Checking daily and stopping on a good day is how you adopt noise.
- Decide on money, not on replies. Reply rate moves easily and often in the wrong direction.
- Accept flat results. Most variables do nothing. Keeping the simpler version is a legitimate outcome and the most frequent one.
- Re-test what you adopt. A result that does not survive a second run on a later period was a season, not a finding.
Then write the answer down next to the sheet, including the failures. An agency that keeps a log of what did not work stops re-litigating the same argument every quarter: the same reading habit as quality control, applied to changes instead of to people.
One last thing worth saying plainly: a test tells you what happened in your inbox, in that window, with those fans. It does not tell you why, and it does not generalise to a creator you have not tested. What our data does not show sets out the same limits for our corpus, and they apply to yours with far less volume behind them.
Frequently asked questions
How long should a chatting test run?
In whole weeks, and until both versions have collected enough sales that your single largest one cannot decide the outcome. Fan behaviour has a weekly shape, so a test that starts on a Tuesday and ends on a Friday has compared different parts of the week. Stopping the moment a version looks ahead is the most common way agencies talk themselves into a change that was never there.
Can I test on one chatter and roll out to the team?
No, and this is the single most expensive mistake in the list. One chatter's results carry their fans, their hours and their skill along with the variable you meant to test. If the version has to run through one person, run both versions through that same person, alternating threads, so their ability sits on both sides of the comparison.
What if I do not have enough conversations to test anything?
Then test sequentially and be honest that it is weaker. Run version A for a full period, version B for the equivalent period, and only accept a difference large enough to be visible without arithmetic. Small inboxes get more from quality control (reading conversations back) than from underpowered experiments.
Should I test on new fans or existing ones?
Pick one and write it down before you start. New subscribers and long-standing fans respond to different things, and a test that lets both into the pool will produce a result that depends on which happened to arrive that week. Most wording tests belong on new subscribers, where the thread history is not doing the work.
Is reply rate a good thing to measure?
It is a useful early signal and a dangerous target. Several phrasings pull more replies while producing fewer sales, because a question invites an answer and the answer can be no. Measure reply rate if you like, decide on revenue, and never let a scorecard reward the first over the second.
What do I do when the test comes back with no difference?
Keep the simpler version and move on. A flat result is a real result: it tells you the variable is not where your money is, and it saves you from defending a change nobody can feel. Most tested variables land here, which is why the ones that do move are worth finding properly.
See what it looks like in practice
The justonedash chatbot holds the conversations, keeps each creator’s voice and works around the clock.