Contact

What is A/B Testing?

Definition

A/B testing is an experiment in which the current version of a page or element (A, the control) and a modified version (B, the variant) are shown at the same time to randomly assigned groups of visitors, to measure statistically which performs better on a predefined metric. To make sure a difference is not down to chance, the sample size is calculated before the test starts and the test runs until it is reached.

Also known as: A/B test, split testing, split test, AB testing, bucket testing

A/B test diagram comparing the conversion rates of two variants with different CTA text on a 50/50 traffic split

What makes it an experiment

Several conditions have to hold together. Visitors are assigned randomly, and each person keeps seeing the same version for the length of the test. Both versions run simultaneously; comparing last month's numbers with this month's variant is a before/after comparison, not an experiment. And one primary metric, for example the share of visitors who submit a quote form, is chosen before launch. Hunting through metrics afterwards for one that "won" is the fastest route to mistaking noise for success.

Multivariate tests, which try combinations of several elements at once, and split URL tests, which serve versions from separate URLs, belong to the same family but need far more traffic.

Sample size comes first

The number of visitors you need depends on four inputs: the baseline conversion rate, the smallest difference worth detecting (the minimum detectable effect), the significance level (typically 5%) and statistical power (typically 80%). With a 3% baseline, detecting a lift to 3.6%, a 20% relative improvement, takes roughly 14,000 visitors per variant under those assumptions. Smaller effects get expensive quickly: halving the detectable lift to 10% roughly quadruples the traffic required.

"Statistically significant" means that a difference this large would be unlikely to appear by chance if the versions truly performed the same, with "unlikely" set by the threshold you chose in advance. It does not tell you how much better the variant is, or whether the difference matters to the business.

Peeking and other traps

  • Peeking. Checking daily and stopping the moment the result looks significant inflates the false-positive rate dramatically. In a classic fixed-horizon test, the duration is set upfront and the result is read only once the sample is complete. If you need to monitor continuously, use a sequential testing method designed for it.
  • Partial weeks. Weekday and weekend behaviour differ, so cover at least one full week, preferably two.
  • Novelty effect. A fresh design can attract extra attention from returning users for a few days, then fade.
  • Sample ratio mismatch. If a 50/50 split arrives as 46/54, assignment or tracking is broken and the result cannot be trusted.
  • Flicker. Client-side testing tools may briefly show version A before swapping in B. That contaminates the test and hurts CLS and loading metrics.

Testing without upsetting search engines

Google's guidance on website testing is explicit. Showing Googlebot different content from users is cloaking and violates its spam policies. If a variant lives on a separate URL, point a canonical URL at the original, and if you redirect, use a temporary 302 rather than a permanent 301. Once the test ends, ship the winner and remove the test code and variant URLs promptly.

Related terms

← Back to the glossary