Running an Agency With 20 Creators: Delegate and Steer
At twenty creators you stop running on feel: how to split portfolios, decide who reviews whom, and steer on spreads between accounts rather than on totals.
Published 28 July 2026
Twenty creators is not four times five. At five, the agency fits in your head and what you do all day is the work itself. At twenty, the work is done by people you are not watching, in inboxes you will never fully read, during hours you are not awake for. You stop reading conversations and start reading numbers about conversations. Everything below exists to make those numbers worth trusting.
What actually changes between five and twenty creators?
You lose direct evidence. At five creators, every problem is something you notice: a fan who sounds confused, a chatter whose messages read differently this week. At twenty, none of those reach you, and the failure signals that worked at five creators stop firing.
What breaks first, in order:
- Reading. You can no longer sample the work yourself, so quality becomes something you are told about rather than something you have seen.
- Memory. Rules that lived in your head (what this creator refuses, what that one discounts) now have to live in documents someone else maintains.
- Averages. With twenty accounts, the agency total goes flat while one account halves and another doubles. The number stops describing anything.
- Escalation. At five, everything reaches you. At twenty, only what someone chooses to escalate reaches you, which means you are steering on a filtered feed.
How do you split twenty creators into portfolios?
Into portfolios balanced on peak hour, with revenue spread across them rather than concentrated. The split you choose determines what your managers can see, so it is worth one deliberate decision rather than a historical accident.
| How you split | What it groups | What it makes easy | What it hides |
|---|---|---|---|
| By revenue rank | Big accounts together, small together | Rewarding your best manager | Small accounts get no attention and never grow |
| By peak hour | Accounts whose fans buy at the same time | Rostering, because one shift covers the whole portfolio | Nothing structural: this is the default to prefer |
| By content type or voice | Similar tone, similar offers | Chatters switching between accounts | Fragile: one bad practice spreads across the whole group |
| By who signed them | The manager who won the account keeps it | Creator relationships | Wildly uneven portfolios, and no comparison between managers |
Peak hour wins because it is the only split that makes the roster cheaper to staff. A portfolio whose accounts peak in the same block can be covered by one shift instead of three, and our data is unambiguous about where that block sits: evenings maximise conversion, and the 2am-6am window collapses it. Balance revenue across the portfolios afterwards, so that comparing two managers means something.
Who reviews whom, and how often?
Review crosses portfolios, always. A manager who scores their own accounts is scoring their own decisions, and the scores drift upward long before the quality does.
A workable chain, weekly:
- The manager reads their own portfolio daily. Unscored, informal, for correction in the moment. This is coaching, not assessment.
- A manager from another portfolio scores a fixed sample. Three threads per chatter: one that went quiet, one that did not close, one that did. Same three every week, so the comparison is stable.
- You read the reviewers, not the chatters. Three scored threads a week, chosen at random from what was already reviewed. You are checking whether the scoring is honest, not whether the writing is good.
- Each creator gets a sample, not the inbox. A handful of threads on a schedule tells her more about the team than the whole inbox, which she will open once.
Score on rules that have been shown to move money, never on taste. That distinction is what quality control means in practice. Publish the grid before anyone’s first shift. A score nobody can see is not correction, and chatters write defensively around it.
What do you keep, and what do you genuinely delegate?
Keep three things, delegate the rest, and mean it. The most common way a twenty-creator agency stalls is an owner who has delegated the work but not the decisions, so every portfolio queues behind one person’s inbox.
| Decision | Who holds it | Why |
|---|---|---|
| Price floor and what is never discounted | You | Reversing a price collapse across twenty accounts takes months |
| Contracts and the commission base | You | The base a percentage applies to recurs every month, in the same direction |
| Escalation list: refunds, disputes, anything legal | You | These carry consequences outside the agency |
| Rota and coverage | Portfolio manager | They know their own peak hours |
| Hiring shortlist and testing | Portfolio manager | The person who lives with the hire should choose it |
| Voice guides and price grids | Manager, written by the creator | A guide written on her behalf produces chatters who sound like the manager |
| Weekly quality sample | Reviewer from another portfolio | Independence is the whole point |
Note what is not on the keep list: chatter selection. Owners hold onto hiring longer than anything else and it is rarely the right call: the protocol matters more than the picker, which is the argument in how to hire chatters.
Which numbers do you steer on at twenty creators?
Spreads, not totals. A flat agency total is the least informative number you own, because it is the sum of movements pointing in opposite directions.
Six lines, read weekly:
- Each creator against her own trailing baseline, never against the roster. Accounts are not comparable to each other; each is comparable to herself a month ago.
- The spread between your best and worst chatter on the same account. Same voice, same price grid, same fans: the gap is a people signal, cleanly isolated.
- The spread between portfolios on the same metric. If one portfolio is consistently better, find out what its manager does differently before you promote them.
- Coverage gap: staffed hours against sales-by-hour. The question is not how many hours you staff, it is whether they are the right ones.
- Share of threads opened. The hours nobody covers appear in no cost line, which is why nobody counts them. That arithmetic is set out in what a chatting team costs.
- Net revenue per creator, on a base written down in words. Gross and net revenue are different numbers, and at twenty accounts the difference is no longer something anyone reconciles by eye.
How do you spot a drop before it costs you?
By watching the conversation, not the revenue line, because a badly handled fan stops replying weeks before he stops appearing in the numbers. Five signals move first, and all five are countable from transcripts.
- Where the pitch lands. Sales attempts drifting earlier in threads is the clearest early warning there is. Our data, across 1.5M messages, puts a pitch before the sixth message at roughly a third of conversion lost against one placed after about ten exchanges. It drifts when chatters are overloaded or paid on activity.
- Reply latency inside the peak block. Slow replies at the hour that converts best cost more than slow replies anywhere else.
- Ellipsis closers reappearing. The worst-performing closer measured, and the easiest habit to slide back into.
- Closed questions at the moment of closing. Costs several points of conversion.
- A personal callback placed at the exact moment of the pitch. It costs points even though it feels like it should help, the most counter-intuitive result in the corpus.
Count each one as a rate per chatter per week. A signal you cannot count is a signal you will argue about.
What does structure not fix?
Two things, and both survive any reorganisation. The first is the hours nobody covers: threads never opened, subscribers nobody has spoken to, lapsed fans nobody has hours for. Splitting twenty accounts into portfolios does not create hours; it only allocates the ones you already pay for.
The second is what a chatter writes at the moment of the pitch. Reviewing, scoring and portfolio structure raise the floor and make drift visible; none of them puts the right sentence in the message. That is a training and tooling question, not an org chart question, and confusing the two is how an agency ends up with three layers of management and the conversion rate it had at five creators.
Where do you start if you restructure one thing this quarter?
Move the scored sample out of the portfolio that produced it. Same people, same grid, different reader. It is the change that makes every other number trustworthy. Until review crosses portfolios, your quality data is a self-assessment, and steering on numbers you cannot trust is worse than steering on feel, because it feels like evidence.
Frequently asked questions
When do you need a reviewer as well as another manager?
The moment the scored sample stops being pulled every week. Another manager buys capacity to run accounts; a reviewer buys the capacity to judge how they are run, and the second one runs out first because nobody chases you for it. Count it directly: threads scored per chatter per week, against threads that should have been. When that number slips two weeks running, you are short a reviewer, not short a manager.
Should portfolios be balanced on revenue or on effort?
On peak hour, with revenue spread across portfolios rather than concentrated in one. A portfolio whose accounts peak in the same block is covered by one shift instead of three. A portfolio of only large accounts looks prestigious and is usually the least improvable, because those accounts are already worked hard.
Can the person who manages an account also review its conversations?
They can read it daily; they should not be the one who scores it. A reviewer marking their own portfolio has an interest in the mark. Keep daily reading with the manager and move the scored sample to someone in another portfolio.
What do you keep as owner once you have twenty creators?
The price floor, the contracts, and the escalation list. Everything else is delegable, and holding onto more than that is how a twenty-creator agency stalls: the owner stays the bottleneck on decisions nobody else is allowed to make. Those three are kept because reversing them later is expensive, not because they are difficult.
Why does the revenue line react so late?
Because a fan who has been handled badly does not stop paying that week, they stop replying. The gap between a conversation going wrong and the money reflecting it is measured in weeks, and it is filled by fans who are still on the list and no longer buying. That lag is the entire argument for steering on conversation-level signals rather than on the monthly total.
Do you need a second layer of management at twenty creators?
Usually one layer plus a reviewer, not a hierarchy. The functions that genuinely need separating are managing accounts, covering shifts, and judging quality. The third one is the one that gets skipped, because it is the only one nobody chases you for. Adding managers without adding a reviewer buys capacity and no visibility.
See what it looks like in practice
The justonedash chatbot holds the conversations, keeps each creator’s voice and works around the clock.