Flat 30% off for Singapore 🇸🇬

A/B test significance & sample size calculator

Find out whether your A/B test result is real, how big the lift could be, and how many visitors you need before you start. P-values, confidence intervals and a Bayesian chance to beat, all in one place.

  • Free, no signup
  • Runs in your browser
  • Nothing is uploaded

Your test

Drag a label sideways to scrub, or type the numbers.

AControl

4.00%

BVariant

4.70%

Confidence level

95% is the usual standard. 90% accepts more false winners; 99% needs more traffic.

Result

Relative uplift (B vs A)

+17.5%

Significant: B wins

p-value 0.0079

needs < 0.05 at 95%

Rate A

4.00%

control

Rate B

4.70%

variant

Difference

0.70

points, B higher

z-score

2.66

two-tailed

Chance B beats A99.6%

Bayesian estimate with flat priors. Read it as a probability, not a guarantee.

B converts at 4.70% against 4.00% for A. At 95% confidence the true difference is between +0.18 pp to +1.22 pp, so B is very likely better. Only stop now if you reached the sample size you planned.

Conversion rates with confidence intervals

Each bar spans the 95% interval; overlap is normal even for real winners.

4.0%5.0%AcontrolBvariant

Difference B − A, 95% interval +0.18 pp to +1.22 pp

B worseB better0

How it works

How to use the a/b test calculator

  1. 01

    Plan the sample first

    On the Sample size tab, enter your baseline rate, the smallest lift worth detecting and your daily traffic. Commit to that number before the test starts.

  2. 02

    Run the test to the end

    Split traffic evenly and leave it running until each variant reaches the planned sample, in whole weeks. Don’t stop early because the result looks good.

  3. 03

    Enter visitors and conversions

    Switch to Test result and type each variant’s visitors and conversions. The p-value, confidence interval and chance to beat update instantly.

  4. 04

    Read the interval, not just the verdict

    The interval shows how big the real difference could be. Share the link so your team sees the same numbers.

How the significance test works

An A/B test compares two conversion rates, each measured on a sample of visitors. Because samples are noisy, B can beat A by chance alone. This calculator runs a two-tailed, two-proportion z-test: it asks how many standard errors apart the two rates are, using the pooled rate of both groups.

z = (pB − pA) / √( p̂ · (1 − p̂) · (1/nA + 1/nB) ), where p̂ is total conversions divided by total visitors. The p-value is the probability of a gap at least this large, in either direction, if the variants were truly identical. Below 0.05 counts as significant at 95% confidence.

The default example on the page shows why this matters: 480 conversions from 12,000 visitors (4.0%) against 564 from 12,000 (4.7%) is a 17.5% relative uplift with p ≈ 0.008. Significant, but the 95% interval for the difference runs from about +0.18 to +1.22 percentage points, so the true lift could be anywhere from modest to large.

Why this calculator is two-tailed

A one-tailed test only looks for an improvement, so it reaches significance with less data. The catch is that it can't tell you when B is significantly worse, and variants that hurt conversion are common. Choosing one tail after seeing which way the data leans is a quiet form of cheating. Two-tailed tests are the safer default for product and marketing experiments, and they are what this calculator uses for both the p-value and the sample size.

Read the confidence interval, not just the verdict

A binary ‘significant / not significant’ hides most of the information. The interval for the absolute difference (calculated with the unpooled standard error) tells you the range of effects consistent with your data. If it is entirely above zero, B is very likely better. If it straddles zero, you haven't shown a difference, which is not the same as showing there isn't one. A narrow interval around zero is a confident ‘no meaningful change’; a wide one just means you need more data.

How the sample size is calculated

Before a test, choose the smallest lift worth detecting (the minimum detectable effect, MDE), the significance level and the power, the chance of detecting the effect if it is real. The calculator uses the standard two-proportion formula:

n = [ z₁₋α/₂ · √(2p̄(1 − p̄)) + z₁₋β · √(p₁(1 − p₁) + p₂(1 − p₂)) ]² / (p₂ − p₁)²

With a 10% baseline, a +20% relative MDE (10% → 12%), α = 0.05 and 80% power, that gives 3,841 visitors per variant. Raise power to 90% and it becomes 5,142. With more than one challenger, the significance level is split across the comparisons (Bonferroni), so each extra variant costs more than its share of traffic.

Baseline rate+10% lift+20% lift+30% lift
1%163,09542,69319,827
2%80,68221,1099,798
5%31,2348,1583,780
10%14,7513,8411,774

Visitors needed per variant at 95% confidence and 80% power. Halving the effect you want to detect roughly quadruples the traffic, which is why small sites should test bold changes, not button colours.

The mistakes that produce false winners

  • Peeking. Checking daily and stopping the moment p drops below 0.05 can push the real false positive rate several times above 5%. Fix the sample size first and evaluate once.
  • Stopping mid-week. Weekday and weekend visitors behave differently. Run whole weeks, at least one.
  • Too many metrics or segments. Test twenty metrics at 95% and you should expect about one false ‘winner’. Pick one primary metric before the test.
  • Sample ratio mismatch. If a 50/50 split delivers 10,000 versus 9,400 visitors, something in the set-up is broken. Fix it before trusting the result.
  • Novelty effects. Returning visitors click new things because they are new. Watch whether the lift fades over the second week.

What the chance to beat adds

The Bayesian ‘chance B beats A’ treats each rate as a Beta distribution updated by your data and estimates the probability that B's true rate is higher. It is easier to explain to stakeholders than a p-value, but it is not a licence to peek either. Use it alongside the interval.

Turning a winner into revenue? Put the lift into the website ROI calculator. And if your tests keep coming back flat, the variants are usually too timid. Bigger, research-led changes are where product design earns its keep.

Common questions

A/B test calculator: FAQ

It means a difference as large as the one you saw would be unlikely if the two versions truly performed the same. At 95% confidence, ‘unlikely’ means a p-value under 0.05. It does not tell you how big the improvement is or guarantee it will hold; the confidence interval answers the size question.

Need it done for you?

Better variants to test start with better design. That’s where I come in.

More free tools

All tools →