What Our Conversation Data Does Not Show: The Limits
The limits of conversation data, stated plainly: what 42,852 conversations and 1.5M messages cannot establish about cause, your creator, or who is missing.
Published 28 July 2026
Our corpus contains 42,852 conversations, 1.5 million messages and 38,879 sales. That is a large dataset by the standards of this industry, and it still cannot answer most of the questions people ask of it.
This page is the list of what it cannot answer: the limits of conversation data, ours included. We publish five measured findings from that corpus and nothing else, and the reason the five are worth reading is that we are willing to draw the line around them.
What is actually in the corpus?
Conversations that happened, and what happened in them. Nothing about intent, nothing about the people, nothing about the business around them.
- The message stream: who wrote, in what order, at what time, how long the thread ran.
- The outcome: whether a sale occurred and where in the thread it sat.
- Not in it: the creator’s contract, the chatter’s pay, the agency’s margin, the fan’s reason for buying or leaving, or anything about a person’s identity.
That last line matters more than it looks. Every question about money, staffing or fan psychology that people put to this dataset is a question about something the dataset does not contain.
Can the data prove that one thing causes another?
No. It is observational, and that word does real work here.
Nobody assigned chatters to conditions. Nobody randomised which fan received which version of a message. Every pattern in the corpus is a comparison between conversations that happened to differ, in inboxes where hundreds of other things differed at the same time.
| The claim | What the corpus supports | What it does not |
|---|---|---|
| Early pitches convert worse | Threads pitched early converted worse, consistently, at scale | That moving your pitch later will lift your number by a stated amount |
| Ellipsis closers underperform | Messages ending that way sold least among endings compared | That the character itself is the cause rather than the writing habit around it |
| Nights convert badly | Conversion collapses in the small hours across the corpus | That staffing those hours differently would recover the sales |
| Callbacks at the pitch cost | Threads doing it converted worse than threads that did not | That the technique is bad everywhere, for every creator |
There is a specific trap behind every row of that table, and it has a name: the third factor. Suppose the chatters who pitch late are also the more experienced ones. The corpus would then show late pitches converting better even if pitch timing did nothing at all, and we would be reading a fact about seniority as a fact about message order. Splitting the population several ways reduces that risk. It never removes it, because the corpus does not record who was typing or how long they had been doing the job.
The distinction is not academic. A causal claim licenses “do this and expect that”. An association licenses “this is where to look first, then test it yourself”, which is the honest instruction and the subject of how to test what works in fan conversations.
What does it say about your creator?
Nothing specific, and this is the limit people ignore most often.
- A population average is not a prediction for one account. Thousands of creators pooled together produce a central tendency. Individual accounts sit across a wide spread around it, and some sit on the wrong side of it for reasons the data cannot see.
- Niche, price tier and audience are not controlled for. A creator selling high-priced custom content to a small group of regulars is a different business from one selling cheap bundles at volume, and both are in there.
- Voice is invisible to the measurement. The thing that most distinguishes one creator’s inbox from another’s is exactly the thing a corpus this size averages away.
Use a finding as a prior, not as a setting. Then check it against your own threads.
Who is missing from the data?
Everyone who never appeared in a conversation, and the gap has a shape rather than being random noise.
- Fans who never wrote. Subscribers who lurked and left generate no thread, so nothing here describes them, and they are a large share of most accounts.
- Creators without agencies. The corpus comes from managed inboxes. Solo creators answering their own messages behave differently and are not represented.
- Threads that were never exported. Deleted accounts, closed platforms, blocked fans. Whatever systematically failed to reach the dataset is systematically absent from it.
- Whole segments of the industry. Languages, regions and price tiers that are thin in the corpus cannot be spoken about, however confident the overall number looks.
A finding is a statement about the observed population. Extending it beyond that population is not analysis; it is optimism.
What does the corpus measure, and over what horizon?
It measures the sale, and the thread around it. It does not measure the year that follows.
- No long-run view of the relationship. A tactic that lifts conversion this week while making a fan tire of the account faster would look, in this data, exactly like a good tactic. Lifetime value is the number that would catch it, and it is not what these findings are built on.
- No measure of how a fan felt. Conversion is the only verdict in the data, and people buy things they are ambivalent about.
- Nothing about what happened to the money afterwards. Refunds, chargebacks, disputes and subscription cancellations after the fact sit outside the measurement.
Why publish the limits at all?
Because a finding you can draw a boundary around is one you can use, and a finding with no boundary is a slogan.
There is also a practical reason. Most numbers circulating in this industry have no source at all: a percentage in a pitch deck, a multiplier in a sales call, a figure repeated until it sounds established. We would rather publish five findings with their edges marked than fifty without, because the corpus is the only asset here that cannot be copied, and one invented figure in it would make the other findings worthless.
So treat every result on this site the same way:
- Read what was compared, and in what population.
- Assume the direction is more reliable than the size.
- Re-run it on your own inbox before you rebuild anything around it.
- Keep the version that survives your own test, not the one that sounded best.
If a page here ever states something the data cannot support, the fault is ours and the finding should be discarded. That is the deal, and it is why the five findings are worth the read.
Frequently asked questions
Does your data prove that pitching early causes lower conversion?
No. It shows that conversations where the pitch landed early converted worse than conversations where it landed later, consistently and across a large number of threads. The most likely explanation is causal, but the corpus is observational, so the honest statement is an association strong enough to act on rather than a proof.
Will these findings hold for my creator?
Probably in direction, not necessarily in size, and we cannot tell you which. A population average is a statement about thousands of accounts pooled together. Your creator has one audience, one price point and one voice, and the only way to know how she sits against the average is to test on her own threads.
Which platforms and countries does the corpus cover?
It covers the accounts whose conversations we have, which is not a random sample of the industry. Any inbox that differs sharply in language, price tier or audience from the ones represented is outside what the measurement can speak to, and findings should be treated as hypotheses there rather than as rules.
Why publish only five findings from a dataset that size?
Because those are the ones that survived. A corpus of this size will produce an endless supply of patterns that look impressive and do not hold up when the population is split, the period changed, or a handful of high-spending fans removed. Publishing the survivors and nothing else is the whole point.
Does the data say anything about how much a chatting team should earn or charge?
Nothing at all. The corpus contains conversations and sales, not payroll, contracts or agency margins. Any figure of that kind you see anywhere is somebody's example or somebody's guess, and it should be presented as such rather than as a measurement.
Could the findings be an artefact of how the data was collected?
It is always possible, and it is the reason each finding is stated with the comparison behind it. Where a pattern could be produced by a collection quirk rather than by fan behaviour, that is a limit of the measurement, and it belongs next to the result instead of in a footnote nobody reads.
See what it looks like in practice
The justonedash chatbot holds the conversations, keeps each creator’s voice and works around the clock.