Conversion & funnel optimisation

A/B testing mistakes to avoid. Where good ideas produce bad data.

The most common A/B testing mistakes are peeking at results before a test finishes, calling a winner from too small a sample, and changing several things at once. Ron Kohavi's research at Microsoft found only about a third of well-designed tests actually improve the target metric, so a disciplined process is what tells a real result apart from noise.

The short answer. Most failed tests fail before the traffic even arrives.

A/B testing mistakes rarely show up as an obviously broken test. They show up as a result that looked clean, got shipped, and then quietly failed to repeat itself a quarter later. Ron Kohavi, who ran experimentation at Microsoft's Bing for years, published research in the Harvard Business Review in 2017 showing that among well-designed, well-powered experiments, roughly a third improve the target metric, a third show no significant difference, and a third actually make it worse. That split is not a sign that testing does not work. It is the reason the process around a test matters as much as the idea being tested, because a two-in-three chance of a flat or negative result is exactly the environment where sloppy process turns noise into a false win.

I see the same handful of mistakes on almost every CRO audit I run, regardless of the industry or the size of the traffic. None of them are exotic. Most come from a reasonable instinct, wanting an answer faster, wanting to test more at once, wanting to trust a dashboard that already looks decisive, and each one is easy to fix once it is named.

The eight mistakes that quietly wreck A/B tests. And the fix for each one.

1. Peeking at results and stopping as soon as they look good

Checking a live results dashboard every day and declaring a winner the first time the p-value crosses significance is the single most common A/B testing mistake, and the most damaging. Statistical significance calculated repeatedly, rather than once at a pre-agreed sample size, inflates the real false-positive rate well past the five per cent a team thinks it has accepted. A test that hits significance on day four, purely by chance, can drift back to flat by day fourteen if you let it run its planned course. The fix is mechanical: decide the sample size and the run length before the test starts, and do not act on the result until both are met.

2. Calling a winner from an underpowered sample

A page converting at two per cent with a few hundred weekly visitors will rarely reach a sample size that can detect anything smaller than a huge lift inside a normal testing window. Running the test anyway and trusting whatever number appears after two weeks produces a result that is really a coin flip dressed up as data. Calculate the minimum sample size from the current conversion rate, the traffic available and the smallest lift worth acting on, before committing a test slot to a low-traffic page at all.

3. Testing several changes at once and learning nothing

A redesigned hero, a new headline and a different button colour shipped together as one variant might beat control, but nobody can say which change did the work. The next page redesign then repeats whichever part was actually irrelevant, because the team is generalising from a bundle rather than a variable. Where a full redesign has to be tested as one unit for practical reasons, follow it with a smaller, single-variable test to isolate what mattered, rather than assuming every element of the winning version earns its place forever.

4. Ignoring seasonality and day-of-week effects

Running a test for five days that happen to include a payday weekend, a public holiday or a marketing campaign spike gives both variants a distorted, non-representative slice of traffic. A test needs to run across at least one full, typical business cycle, in most B2B and ecommerce contexts that means a minimum of two full weeks, so weekday and weekend behaviour lands in both variants roughly equally rather than skewing whichever one happened to catch the unusual traffic.

5. Averaging away a result that only holds for one segment

An aggregate result can hide the opposite pattern happening underneath it. A new checkout flow might lift conversion sharply for returning customers while quietly hurting first-time visitors, and if first-time traffic is the smaller slice, the blended number still reads as a win. Segment the result by new versus returning, and by device, before shipping anything, because a change that only works for part of the audience needs a different rollout decision than one that works for everyone.

6. Testing ideas with no hypothesis behind them

Running a test because a competitor did something similar, or because someone has a hunch, fills the testing calendar without building any real understanding of why visitors behave the way they do. A test built from a specific hypothesis, rooted in a heatmap finding, a piece of user research or a support ticket pattern, teaches something whether it wins or loses. A test built from a guess only teaches something if it happens to win, and most will not.

7. Shipping a variant with a tracking or implementation bug

A flicker effect where visitors briefly see the control before the variant loads, a conversion event that fires twice on one variant, or a mobile layout that breaks only in the test version, all of these quietly corrupt a result without ever throwing an error a team would notice. QA both variants on every major device and browser before launch, and check that the primary conversion event and any secondary metrics fire exactly once per real action, not zero or twice.

8. Treating one win as a permanent truth

A result that clears significance once is still a sample, not a law. Regression to the mean means a genuinely strong first result is somewhat likely to look weaker on a retest, and a change that wins in one season, one traffic mix or one campaign context can lose once those conditions shift. Big, high-stakes wins, a pricing page change or a new checkout flow, are worth a confirmation retest before they get treated as permanent, especially several months on from the original test.

A worked example. Where a pricing page test went wrong, twice.

A SaaS client came to me convinced their pricing page redesign had failed, because a follow-up test on the same page showed no lift, months after an earlier test had shown a clear seven per cent improvement. Digging into the first test's setup explained the gap: it had run for nine days, been checked daily, and been called a winner the moment it crossed significance on day six, which happened to fall during a week the marketing team was running a paid campaign that skewed heavily toward one segment. The second test, run properly for a full three weeks with the sample size calculated up front, showed the redesign was roughly neutral overall, with a real but modest lift for mobile visitors specifically and a small loss for desktop. Neither test was dishonest. The first one was just answering a different, narrower question than the team thought it was asking, and calling it a general win overstated what it had actually shown.

That is the pattern behind most A/B testing mistakes: not bad intentions, but a process loose enough to let a real but narrow result get reported as a bigger and more permanent one than the data supports. Our work on conversion rate optimisation starts by tightening exactly that process, sample size, run length and segmentation agreed before a single visitor sees a variant, because the fix is cheaper before a test launches than after a false winner has already shaped the roadmap.

A pre-launch checklist. Five questions to answer before a test goes live.

  • What is the hypothesis, in one sentence? Rooted in a specific piece of evidence, not a preference.
  • What sample size and run length end the test? Calculated in advance, from current traffic and conversion rate.
  • Is exactly one meaningful variable changing? Or is this a bundled test that needs a single-variable follow-up.
  • Have both variants been QA'd on the devices real traffic actually uses? Not just the one on the designer's desk.
  • Which segments will the result get checked against? New versus returning, and device, at minimum, before the result gets generalised.

None of this replaces judgement, and a small business without the traffic for classic significance testing still has good options, heatmaps and session recordings read a page's problems without needing thousands of visitors first. What the checklist does is make sure that whichever method is used, the result that reaches the roadmap is the one the data actually supports, not the one the team was hoping to see. Our broader guide to calculating conversion rate covers the maths underneath the significance calculation, for anyone building this into a testing programme from scratch.

Common questions.

What is the single biggest A/B testing mistake?

Peeking at results before the test has run its planned duration and stopping the moment the number looks good. Checking a live dashboard repeatedly and calling a winner the first time it crosses significance inflates the false-positive rate far above the five per cent a team thinks it is accepting, because each look is another chance for random noise to look like a real effect.

How long should an A/B test run before you trust the result?

Long enough to hit the sample size calculated before the test started, and at least one full business cycle, usually two full weeks, so weekday and weekend behaviour and any day-of-week effect are represented in both variants equally. Stopping early because a dashboard looks promising is how a temporary swing gets mistaken for a real result.

Can you test more than one change on a page at the same time?

Only if you are prepared not to know which change caused the result. A test that changes the headline, the button colour and the layout together might win, but a single-variable follow-up test is the only way to learn which change actually did the work, and that follow-up test is usually skipped once the first one looks successful.

What sample size do you need for a reliable A/B test?

Enough to detect the smallest lift worth acting on, calculated before the test starts from your current conversion rate, traffic volume and the minimum effect size that would justify shipping the change. A page converting at two per cent with a few hundred visitors a week will rarely reach a trustworthy sample inside a normal testing window, which rules out testing it as often as the roadmap wants to.

Why did my A/B test show a winner that did not hold up later?

Usually one of three reasons: the test was stopped as soon as it looked significant rather than at its planned end, the win was true for one segment, new visitors or mobile traffic, for example, but got averaged away across everyone, or a novelty effect from the change wore off once regular visitors stopped noticing it was different. Ron Kohavi's research at Microsoft found roughly a third of well-designed tests improve the target metric, a third are flat and a third are negative, so a result that does not replicate is the norm to plan for, not the exception.

Not sure whether your last test result would survive a retest?

Book a short call and we'll review your testing process, sample sizes and segmentation, so the wins you ship are the ones that actually hold up.

Let's talk ↑