Strategy 6 min read

How Much Traffic Do You Need to A/B Test?

The uncomfortable math of sample sizes, why 30k monthly sessions is our floor, and what to do instead if you are not there yet.

UIUX ServicePublished

The most common question we get from Shopify founders is not about tactics. It is "can we even A/B test yet?" It deserves an honest answer, because running underpowered tests is worse than running none: you get confident-looking numbers that are actually coin flips.

The math, without the jargon

To detect a 10% relative improvement on a 3% conversion rate with standard confidence settings, you need roughly 50,000 visitors per variant, 100,000 total. Chasing a smaller, more realistic 5% lift quadruples that. This is why so many "winners" from low-traffic tests vanish after launch: the test never had the statistical power to see what it claimed to see.

Why we set the floor at 30k sessions/month

With 30k+ monthly sessions, we can point tests at high-traffic, high-intent pages (PDP, cart) and target the larger effect sizes that good UX engineering produces, typically reaching 95% confidence within a 2–4 week window. Below that, test durations stretch past six weeks, seasonality pollutes the sample, and the testing program stalls.

Under 30k sessions? Do this instead

  • Evidence-led redesign: fix the leaks that session recordings, GA4 funnels and heuristic review agree on. You do not need a test to prove a broken variant selector is broken.
  • Test bigger swings: radical page variants, not button colors. Larger true effects need far smaller samples.
  • Measure sequentially with guardrails: before/after with GA4 cohort comparison is imperfect but honest when clearly labeled.
  • Bank the traffic math: put the effort into SEO and retention that gets you to testable volume sooner.
A CRO program is not "run tests." It is "make evidence-based improvements at the fastest speed your traffic statistically allows." Sometimes that speed is a redesign, not a split test.

The PIE framework we score every idea with

Potential (how big is the lift if it works), Importance (how much traffic and revenue flows through the page), Ease (how cheap is it to build and test). Every hypothesis in our 90-day backlog carries a PIE score, so the first test we ship is the one with the best expected value, not the one that was easiest to imagine.