Validate decisions

A/B significance calculator

You get both rates, their difference, the p-value when the sample supports it, and a significant or inconclusive conclusion.

Free · no sign-up Estimated time: 3 minutes This calculation runs on your device. We do not save your inputs.
Direct answer

Is the difference between A and B statistically significant?

The calculator runs a two-sided Z-test for two independent proportions, using a pooled rate under the null hypothesis that A and B are equal. The p-value shows how incompatible a difference this extreme or more would be with that hypothesis. A conclusion is shown only when the expected success and failure counts in both groups are at least 5.

A/B test example

A converts 100 of 1,000 participants and B converts 130 of 1,000: 10% versus 13%, a +3-point absolute difference and +30% relative lift. The test gives z≈2.10 and p≈0.0355; at α=0.05, the difference is significant and favors B.

Your inputs

Add your numbers or start with sample data.

95% is the most common choice
This calculation runs on your device. We do not save your inputs.

Your result

Complete the fields to see the result and interpretation here.

Practical lab

Understand it in a few minutes

  1. 01

    Statistical significance does not tell you whether a gain matters to the business.

  2. 02

    A two-sided test is safer when the variant could also make the result worse.

  3. 03

    Too little data can hide a real effect or produce an unstable conclusion.

How we reached this result

We calculate z = (rate B − rate A) ÷ pooled standard error, then obtain the p-value from the normal distribution.

What this calculation does not show

The test assumes independent groups and a stable experiment; multiple comparisons need extra care.

Direct answer

Frequently asked questions

Which statistical test does the calculator use?

It uses a two-sided Z-test for two independent proportions, with standard error based on the pooled proportion under the null hypothesis of equal rates.

Why must every expected count be at least 5?

The test relies on a normal approximation. Smaller expected counts make it unstable, so the tool asks for more data instead of naming a winner.

Does p<0.05 prove that B is better?

No. It indicates incompatibility with equal rates under the test assumptions. Randomization, experiment quality, effect size, and business impact still require evaluation.

Can I stop at the first significant result?

Avoid repeated peeking and unplanned stopping because they increase false positives. Set the sample, duration, metric, and decision policy before the experiment starts.

Methodology

Sources, assumptions, and review

The references support the definitions and assumptions. The formula used is shown on this page and covered by automated tests.

Tools and learning

If the result is inconclusive, continue to the planned sample instead of stopping at the first positive sign.

Put the learning into a real survey and track every response in NPSLab.

Create a free survey