Rankite
ServicesResultsToolsTeamAboutBlogCareersContactFree SEO Audit
Free tool

Sample Ratio Mismatch Calculator

Enter your observed visitor counts and your planned split to check your A/B test for SRM, instantly and free.

Home / Tools / Sample Ratio Mismatch Calculator
Chi-square statistic
-
P-value
-
Expected A / B
-
Actual split
-
Verdict
-

Enter your observed and planned visitor counts to run the chi-square test.

Built by Rankite, the SEO team behind Swordfish AI's +400% revenue and Zluri's +45% organic growth. See the case studies

Sample ratio mismatch, or SRM, is what happens when the actual number of visitors in each arm of your A/B test does not match the split you configured. Set up a 50/50 test and end up with 10,245 users in variant A against 9,758 in variant B, and the question becomes whether that gap is normal noise or a sign your experiment is broken. The calculator above runs a chi-square goodness-of-fit test on your numbers and gives you a p-value, the standard way experimentation teams answer that question before they trust a test result.

How SRM is calculated

The test compares your observed counts against the counts you would expect if the split had gone exactly as planned. For each variant, you take the squared difference between observed and expected, divide by the expected count, then sum across variants to get a chi-square statistic. That statistic converts into a p-value, the probability of seeing an imbalance this large or larger purely by chance if your randomization were working correctly. A very small p-value means the imbalance is very unlikely to be chance.

What p-value means SRM

Most experimentation teams use a stricter bar for SRM than the 0.05 cutoff common in significance testing. Because an SRM check runs on every experiment you launch, a 0.05 threshold would flag false alarms constantly just from normal variation. The convention that has become standard across the industry, popularized by Microsoft's experimentation platform team, is to flag SRM at p below 0.001, with the 0.001 to 0.01 range treated as a borderline zone worth a second look rather than an automatic fail.

Why sample ratio mismatch matters

When SRM is present, the two groups you are comparing are no longer the same kind of user in aggregate, which breaks the core assumption behind an A/B test. A skewed split is often a symptom of something upstream: bot traffic hitting one variant unevenly, a redirect or caching layer leaking users before they are logged, broken randomization in the assignment code, or one variant loading slower and losing impatient visitors before the experiment records their exposure. Any of these can quietly bias your results in the same direction as, or opposite to, the effect you are trying to measure, which is how a broken test ships a confidently wrong decision. If your team is running experiments on landing pages or funnels and wants that program built on a solid measurement foundation, a free SEO audit is a good place to start the conversation.

Related articles

FAQ

Sample Ratio Mismatch Calculator: questions, answered

What is sample ratio mismatch (SRM)?
Sample ratio mismatch happens when the number of visitors actually assigned to each variant of an A/B test does not match the split you configured. If you set up a 50/50 test but end up with 55,000 users in variant A and 44,000 in variant B, that gap is bigger than random chance would produce, and it is a sign something in the experiment setup or measurement is broken.
How is SRM calculated?
SRM is tested with a chi-square goodness-of-fit test. You compare the observed visitor counts in each variant against the counts you would expect from your configured split, sum the squared difference divided by the expected count for each variant, and convert that chi-square statistic into a p-value. A very low p-value means the imbalance is unlikely to be random.
What p-value indicates SRM?
Most experimentation teams flag SRM at a p-value below 0.001, a much stricter bar than the usual 0.05 used for test results, because an SRM check runs on every single experiment and a looser threshold would trigger constant false alarms. A p-value between 0.001 and 0.01 is usually treated as a borderline warning worth a second look.
Why does SRM matter?
When SRM is present, the two groups in your test are no longer comparable, which means any lift or drop you measured could be caused by the imbalance itself rather than the change you tested. Shipping a decision based on a test with SRM is a common way flawed A/B tests produce confidently wrong answers.
What causes sample ratio mismatch?
Common causes include bot traffic hitting one variant more than another, a redirect or caching layer that leaks users out of the test before they are counted, broken randomization in the assignment code, or one variant loading slower and losing users to bounces before the experiment logs the exposure.

More free tools

Let's grow

Ready to own page one?

Get a free, no-obligation SEO audit and a 30-minute strategy session. We'll show you exactly where the growth is hiding.

Book your free audit Explore services
Get in touch

Tell us about your project

Fill out the form and we'll get back to you within one business day. Prefer email? Write to us directly at contact@rankite.com.

Or copy our email and write to us directly: contact@rankite.com