What Is A/B Testing and How Do You Run One?
A controlled experiment that shows two variants to real users at the same time to measure which one performs better.
Book a free callA/B testing (split testing) is a method of testing two versions of a page, creative or flow at the same time by randomly splitting users into two groups. The goal is to determine with statistical confidence which variant performs better on the target metric (conversions, clicks, revenue). It lets you make decisions based on evidence rather than intuition.
How do you set up an A/B test?
- Write a hypothesis. "If we use benefit-driven copy in our screenshots, store conversion will go up, because users will understand what the app does faster." A hypothesis without a "because" keeps you from learning from the result.
- Pick one variable. If you change several things at once, you won't know which one worked.
- Calculate sample size upfront. Before launching, determine how many users and how much time you need.
- Split traffic randomly. Make sure segments are evenly distributed.
- Run it to completion. Don't stop the test before you reach the target sample.
- Document the result. Losing tests are as much a part of your institutional knowledge as winning ones.
What does statistical significance mean?
It's the likelihood that the difference you see between two variants isn't due to chance. The common threshold is a 95% confidence level, meaning there's less than a 5% probability that the result is random.
Significance alone isn't enough. Evaluate three things together: effect size (how big the difference is), confidence level (how sure we are) and practical significance (whether the difference creates business value). A 0.3% improvement confirmed at 95% confidence may not be worth the cost of implementing it.
How long should a test run?
- At least one full week: Weekday and weekend behavior differ, and short tests miss that cycle.
- Until you reach the target sample: Accumulated data, not elapsed time, is what decides.
- No longer than two weeks: In drawn-out tests, outside factors (campaigns, seasonality, product changes) contaminate the result.
- Longer for lagging metrics: If you're testing retention or subscription renewals, the measurement window naturally gets longer.
Common mistakes
- Stopping early. Declaring the variant that leads in the first few days the winner is the most common mistake. With small samples the difference is highly volatile and often reverses.
- Too many variables at once. You'll know which variant won but not why, so learning doesn't accumulate.
- Not enough traffic. On low-volume pages, the time needed for a significant result may be impractical; in that case, test bigger changes.
- Optimizing the wrong metric. A change that lifts clicks can lower purchases. Always track the down-funnel metric too.
- Not documenting results. Retesting an idea months later that has already been tested is a serious waste of time.
- Overlapping tests. Tests that overlap on the same user group muddy each other's results.
Where can you A/B test on mobile?
- Google Play: Store Listing Experiments let you test icons, screenshots and descriptions natively.
- App Store: Product Page Optimization supports testing up to three variants.
- Ad creatives: On Meta, TikTok and Google, creative variant tests run within the campaign.
- In-app: Onboarding flows, paywall design and pricing tests run on feature-flag infrastructure.
- Landing pages: Headline, visual and CTA tests on the web.
Frequently asked questions
How much traffic do you need?
It depends on your current conversion rate and the size of the improvement you want to detect. Detecting small improvements requires a much larger sample. For low-traffic products, the right strategy is to test big changes that create clear differences rather than small tweaks.
What's the difference between A/B testing and multivariate testing?
An A/B test compares two variants and isolates a single variable. A multivariate test tries several elements at once in different combinations; it tells you more but requires much more traffic. For most teams, A/B testing is more practical.
Is a losing test a failure?
No. Learning that a hypothesis was wrong is a valuable result that keeps you from investing in the wrong direction. It's normal for a large share of tests to show no significant difference; the real failure is moving forward on assumptions without testing.
Let's build a testing culture together
We set up a disciplined testing loop, from hypothesis to measurement, so your decisions are grounded in data.
Book a free call