A/B testing works by splitting your audience into two groups and showing each group a different version of the same element, then measuring which version performs better against a defined goal. One group sees the original (the control), the other sees the variation, and the results tell you which version drives more of the behaviour you want. The sections below unpack every key question marketers have about running effective A/B tests.
What gets changed in an A/B test?
In an A/B test, you change a single variable between the two versions so you can isolate its effect on performance. Common elements include subject lines, call-to-action button text, headlines, images, email send times, landing page layouts, and form lengths. Changing only one thing at a time is the golden rule, because it ensures any difference in results can be attributed to that specific change.
The most productive variables to test tend to be those closest to the conversion action. A subject line change affects whether your email gets opened at all, so it sits at the top of the funnel. A button colour or copy change affects whether someone who is already engaged takes the next step. Both are worth testing, but the closer the variable is to the goal, the more directly its impact shows up in your data.
How does an A/B test actually run?
An A/B test runs by randomly dividing your audience into two equal segments, exposing each segment to one version of the element being tested, and then tracking how each group responds over a set period. Randomisation is essential because it removes selection bias, ensuring the two groups are comparable before the test begins.
Most marketing platforms handle the split automatically. You define the variable, set the audience size, choose the duration, and let the tool distribute traffic or recipients. Throughout the test, both versions run simultaneously under the same conditions, which controls for external factors like the day of the week or seasonal behaviour. Once the test period ends, the platform reports the results and you decide whether to adopt the winning version.
What is statistical significance in A/B testing?
Statistical significance in A/B testing is a measure of how confident you can be that the difference in results between two versions is real and not due to random chance. It is typically expressed as a confidence level, with 95% being the standard threshold most marketers use. If your result reaches 95% confidence, there is only a 5% probability that the observed difference happened by chance.
Reaching statistical significance requires a large enough sample size and enough conversions in each group. A test that shows one version outperforming another by a wide margin on a tiny sample is not reliable. Running a test before you have enough data is one of the most common mistakes in A/B testing, and it leads to decisions based on noise rather than genuine insight.
How long should an A/B test run?
An A/B test should run long enough to collect a statistically significant sample and to account for natural variation in audience behaviour across different days and times. For most email campaigns, this means waiting until the full send has been received and engaged with. For website tests, the general guidance is to run for at least one to two full business cycles, which is typically one to two weeks minimum.
Ending a test early because one version looks like it is winning is a common mistake. Early results are often misleading because the first people to engage with content are not always representative of your broader audience. Patience is a genuine competitive advantage in A/B testing. If your traffic or send volume is low, you may need to run a test for longer to accumulate enough data for the result to be trustworthy.
What’s the difference between A/B testing and multivariate testing?
A/B testing compares two versions of a single variable, while multivariate testing simultaneously tests multiple variables and their combinations to identify which combination of changes produces the best result. A/B testing is simpler, faster, and requires a smaller audience. Multivariate testing is more complex and requires significantly more traffic to produce reliable results.
For example, an A/B test might compare two different subject lines. A multivariate test might simultaneously test two subject lines, two sender names, and two preview texts, generating multiple combinations and measuring which combination performs best overall. Multivariate testing is most useful when you want to understand how elements interact with each other, rather than just which single change has the most impact. For most teams, A/B testing is the more practical starting point.
Which metrics should you measure in an A/B test?
The metrics you measure in an A/B test should be directly tied to the goal of the element you are testing. There is a primary metric, which is the one your test is designed to move, and secondary metrics, which help you understand the broader impact of the change.
Primary metrics by test type
- Email subject line: open rate
- Call-to-action button: click-through rate
- Landing page layout: conversion rate or form submission rate
- Send time: open rate or revenue per email
Secondary metrics to watch
- Unsubscribe rate: a version that lifts opens but also drives unsubscribes is not a net win
- Revenue per recipient: especially relevant for e-commerce campaigns
- Bounce rate: useful for landing page tests to assess engagement quality
- Time on page: helps distinguish between clicks driven by curiosity versus genuine intent
Focusing only on the primary metric without checking secondary metrics can lead you to declare a winner that is actually causing harm elsewhere in the funnel.
Why do A/B test results sometimes fail to hold up?
A/B test results fail to hold up when the winning version is applied to a broader audience or different context and no longer performs as expected. This is sometimes called the novelty effect, where a new variation gets a temporary lift simply because it is different, not because it is genuinely better. It can also happen because the test audience was not representative of the full audience.
Other common reasons results do not hold include running the test during an unusual period such as a promotional event or a holiday, stopping the test too early before the data stabilised, or testing on too small a segment. External factors, such as a competitor campaign or a news event, can also skew results in ways that are hard to detect during the test itself. When a result does not replicate, it is worth investigating whether the original test conditions were truly controlled and representative.
What should you do after an A/B test ends?
After an A/B test ends, you should document the result, implement the winning version if the result is statistically significant, and then plan your next test. A single A/B test is the beginning of an optimisation process, not the end of one. The insight from each test should inform what you test next.
Documenting results is often overlooked but is genuinely valuable. Recording what was tested, what the hypothesis was, what the result was, and what you learned builds an institutional knowledge base that prevents teams from repeating tests that have already been run. It also helps new team members understand what has already been optimised and what assumptions have been validated. Over time, a well-maintained testing log becomes one of the most useful assets a marketing team can have.
How Spotler helps with A/B testing
We built A/B testing directly into our platform so that running structured experiments does not require a separate tool or technical setup. Within Spotler, you can test and optimise across multiple channels and touchpoints as part of your existing marketing workflow.
- Email A/B testing: Test subject lines, sender names, content blocks, and send times directly within your email campaigns, with automatic winner selection based on your chosen metric.
- Website personalisation testing: Our Website Personalisation tool includes built-in A/B testing so you can measure which personalised content blocks, overlays, or page variations perform best for specific audience segments.
- Audience segmentation: Our CDP builds enriched visitor and contact profiles that make it easier to test across meaningful segments rather than your full database, giving you more precise and actionable results.
- Cross-channel insight: Because our tools are connected within the Spotler Marketing Cloud, the insights from one test can feed directly into segmentation and personalisation across email, website, and other channels.
If you want to start running smarter A/B tests without adding complexity to your stack, get in touch with our team to see how Spotler fits your workflow.
Frequently Asked Questions
How do I know what sample size I need before starting an A/B test?
Before launching a test, use a sample size calculator (many are available free online) and input your current baseline conversion rate, the minimum improvement you want to detect, and your desired confidence level. As a rough guide, most email tests need at least 1,000 recipients per variant, while website tests often require several thousand sessions per variant to produce reliable results. Running a sample size calculation upfront prevents you from ending a test too early or wasting time on a test that your traffic volume cannot support.
Can I run more than one A/B test at the same time?
Yes, but only if the tests are running on entirely separate audience segments or different channels, so that one test cannot influence the results of another. If you run two simultaneous tests on overlapping audiences — for example, testing both a subject line and a landing page for the same campaign — you lose the ability to attribute results cleanly to a single variable. The safest approach is to sequence tests on the same funnel stage rather than stacking them, and to use your platform's audience segmentation tools to keep test groups isolated.
What should I do if neither version produces a clear winner?
A null result — where neither version significantly outperforms the other — is still a valid and useful outcome. It tells you that the variable you tested does not have a meaningful impact on performance, which frees you to focus your testing efforts elsewhere. Before moving on, check that the test ran long enough, had sufficient sample size, and that your hypothesis was specific enough to produce a detectable difference. If everything checks out, log the result and shift your attention to a variable that is more likely to move the needle.
How do I prioritise which elements to test first?
A useful framework is to prioritise tests by the combination of potential impact, ease of implementation, and how close the element is to your primary conversion goal. Elements that appear early in the user journey — such as subject lines or hero headlines — affect every subsequent step, so improvements there compound across the funnel. Start with the highest-traffic or highest-volume touchpoints in your workflow, as these will reach statistical significance fastest and deliver learnings you can act on sooner.
Is it a problem if my A/B test audience behaves differently at weekends versus weekdays?
Yes, this is a real consideration, which is why the guidance to run tests for at least one to two full business cycles exists. If your audience skews towards weekday engagement, a test that only captures weekend behaviour will not be representative. For email campaigns, ensure your send captures your typical open window in full. For website tests, always include at least one complete weekend in your test window so that behavioural variation across the week is evenly distributed across both variants.
How do I avoid the novelty effect skewing my A/B test results?
The novelty effect — where a new variation gets a temporary lift simply because it is unfamiliar — is most common in website and interface tests rather than email tests, since email recipients rarely remember what a previous version looked like. To mitigate it, run your test long enough for initial curiosity-driven behaviour to settle, and pay close attention to whether the performance gap between variants narrows over time. If the winning variant's lead shrinks significantly as the test progresses, treat the result with caution and consider extending the test duration before making a permanent change.
What is the best way to build a long-term A/B testing programme rather than running one-off tests?
The foundation of a sustainable testing programme is a structured backlog of hypotheses, a consistent process for documenting results, and a regular cadence for reviewing what has been learned. Start by identifying the key conversion points across your marketing funnel and generating at least one testable hypothesis for each. After each test concludes, hold a brief review to extract the insight, update your testing log, and use the finding to generate the next hypothesis. Over time, this iterative loop builds a compounding body of knowledge about your audience that becomes one of your most valuable marketing assets.