
How to Prioritize Growth Experiments Without Burning Out
- Rhys Mohun
- Jul 6
- 4 min read
How to Prioritize Growth Tests Without Turning Your Backlog Into a Popularity Contest
Boasting a massive test backlog isn't always a sign of a successful experimentation program. Often a team has more opinions than discipline, and rarely a method for how to prioritize growth tests for humans when everything feels urgent.
In my work with growth teams across Canada and the US, the problem I see is a dichotomy: a shortage of ideas, or, too many ideas that spark debate and disagreement. A team has to agree on how to prioritize experiments - not an individual.
How we mess this up
Most teams will say to prioritize by impact. Sean Ellis' classic ICE framework will tell us to line up the home runs at the top. Go big or go home. This works for a while.
Then we remember we're a team of humans, not robots, and we rarely keep up with the velocity predicted in the velocity delivered. In short - we burn out.
ICE had a good run, but the model hasn't evolved with the way today's generation works. Here's what a pure ICE prioritization framework still gets wrong:
The sum of scores. Adding the 3 scores together and ranking by total - this opens us up to imbalanced work. What we actually want isn't a ranked list. It's a balanced backlog: a mix of bets that reflects where the team has capacity, where it needs a win, and where it needs energy back. A sum can't do that human math for us.
Using vague or subjective 1-5 ratings. A "4 for impact" means something different to everyone in the room unless the scale is tied to something concrete. I'll go deeper on what that looks like another time. For now: if a rating scale doesn't force people to point at evidence, it's just an opinion acting as data.
Remove the subjectivity we can, name the subjectivity we can't
I don't advocate for a complicated scorecard. Instead it's agreeing on what makes a test worth running, in that moment, and being honest about where subjectivity lies. Subjectivity is impossible to completely remove within an organization, so we embrace it where necessary.
I score every test against four dimensions I call the 4Es: Evidence, Effect, Effort, Excitement. Evidence and Effect strip out guesswork by tying the score to something observable rather than a gut feeling. Use numbers and data here. Effort forces an honest look at what a test actually costs in build time, whatever that looks like in your org. Often I like to use "heads involved" as a measure where we can't get as accurate as an epic or story size. Excitement is the one I learned to add at Intuit, and it's responsible for some of the best career highlights in my memory :)
None of this makes the process objective. It makes the subjective part visible and defensible instead of buried in "I just have a good feeling about this one." We're not robots. Prioritizing growth tests still needs a human in the room weighing the plate, not a formula outputting a rank.
A balanced backlog has a strategy. An ICE ranked backlog doesn't.
All else equal, effect does come first. But all else is rarely equal, and treating your test backlog like a spreadsheet completely misses the point of prioritization.
A few situations where I'll deliberately deviate from a purely impact-driven plate of work:
The team is coming off a drought of losses. This is the moment to lean into higher-evidence tests, not because they're safer for their own sake, but because leadership, the board, or partners need to see confidence rebuilt before they'll back a bigger swing. Optics are real, like it or not.
Capacity is genuinely constrained. A big project is coming, or it's vacation season and the team can't deliver at normal volume. The right move is a lower-effort plate for that cycle, not pretending capacity is infinite.
The team feels unheard. This one gets ignored the most and it shouldn't. If alignment is shaky and morale is low, deliberately weighting toward tests the team is excited about rebuilds buy-in in a way no all-hands meeting will. This has mattered more for team performance in my experience than almost anything else on this list.
A balanced backlog accounts for burnout, confidence, and business context. A ranked backlog just accounts for the biggest number.
Sorting the backlog on quality not quantity
When I review a test backlog, I don't treat every idea as equally ready. I sort into three groups. I use quality gates to ensure nothing hits the work queue without the fundamentals.
Ready to load. Clear hypothesis, measurable outcome, enough evidence, a realistic delivery path. This test can be safely entered into the queue.
Needs revisit. The opportunity is real, but the hypothesis is vague or the implementation path is muddy. Good ideas with weak framing or evidence. Before forcing our leaders or stakeholders to tell us - let's revisit this one ourselves.
Parking lot. Interesting, possible, not ready. Visible without stealing focus from what is. A parking lot is a viable, no shame place to hang out and get some air. Plenty of great ideas still live here - but they're in a holding pattern.
This alone changes the tone of planning. You're not choosing from forty-seven competing notions. You're choosing from a short list of tests that have earned the right to be considered.
The silly question that cuts through the noise
Before I promote a middling test, I ask one thing: what would need to be true for our customers to want to click this? Did we say or create something worth clicking?
There are also a slew of great "check for boldness" type exercises you can do with your teams, which I'll cover in another post.
If the answer is vague, the test is too. If it reveals three risky assumptions stacked on top of each other, confidence should drop. If it points to a piece of buyer behavior you don't yet understand, the learning value probably just went up.
Prioritizing growth tests isn't about scoring faster. It's about thinking better, and being honest about which of those four dimensions is actually driving the decision this cycle.
If you take one thing from this: stop summing. Start balancing. A test backlog built on a shared, visible logic behaves like an asset. One built on a leaderboard just becomes a longer list of things nobody agrees on.




Comments