🧪

A/B Test Significance Calculator

Enter visitor and conversion counts for both variants to calculate statistical significance, p-value, confidence level, and whether you have a clear winner.

Control (A)

Conversion Rate

Variant (B)

Conversion Rate
Confidence Level
90%
95%
Relative Uplift
B vs A conversion rate
Z-Score
Standard deviations from mean
P-Value (two-tailed)
Probability result is due to chance
95% Confidence Intervals
Control (A)
Variant (B)
Estimated Sample Size Needed (per variant)
To reach 95% confidence detecting the current observed uplift
Advertisement
About this A/B test calculator

This A/B test significance calculator tells you whether the difference between two versions of a page, email, or feature is a real effect or just random noise.

Enter the visitors and conversions for your control and your variant, and it instantly returns the relative uplift, z-score, two-tailed p-value, confidence level, and 95% confidence intervals for each rate. A plain-English verdict tells you whether you can safely declare a winner yet. People use it to:

All of the statistics run locally in your browser, so your visitor and conversion numbers never leave your device — no account, upload, or tracking involved.

Rather than guessing from raw conversion rates, the calculator applies a proper two-proportion z-test so you avoid the most common mistake in optimization: calling a winner too early on a difference that would vanish with more data. Bookmark it as a quick gut-check for every experiment you run.

How to use
  1. Enter visitor and conversion counts for your control (A) — the original, unchanged baseline.
  2. Enter visitor and conversion counts for your variant (B), the new version you are testing.
  3. Watch the conversion rates, relative uplift, and confidence meter update live as you type.
  4. Read the verdict and p-value to see whether the difference is statistically significant.
  5. If you are not yet significant, check the estimated sample size to know how much more traffic you need before declaring a winner.
FAQ

A two-tailed z-test for two proportions with a pooled standard error, plus 95% Wald confidence intervals for each variant. The sample-size estimate targets 95% confidence at 80% power.

95% (p < 0.05) is the industry default. Use 99% for high-stakes changes like pricing or checkout. Decide your threshold before launching the test — never after seeing the results.

Usually no. 94% means a 6% chance the result is noise, above the standard threshold. Wait for more data unless traffic is severely limited and the business risk is low.

There is no single magic number — it depends on your baseline conversion rate and the size of the lift you want to detect. Smaller effects and lower conversion rates need far more traffic. As a rough guide, many tests need a few thousand visitors per variant, but the estimated sample size shown above is tailored to your own numbers.

It means the difference you measured is unlikely to be the result of random chance. At 95% confidence, there is only about a 5% probability you would see a gap this large if the two versions were truly identical. It does not promise the variant is better — only that the data is strong enough to act on.

95% is the standard for most marketing and product tests and balances speed against certainty. Choose 99% when a wrong call is expensive, such as pricing, checkout, or anything affecting revenue at scale. Higher confidence demands more traffic and patience, so match the threshold to the stakes of the decision.

Most often you simply do not have enough data — the difference between variants is small relative to the traffic collected, so the noise still outweighs the signal. Let the test run longer, ideally covering full weekly cycles, and avoid peeking and stopping the moment it briefly crosses the line. The sample size estimate tells you roughly how much further you have to go.