At a glance
Test a justified change under conditions that are as comparable as possible. Select the target metric and evaluation time point before starting, and test both variants technically. Minor differences in results or small sample sizes do not automatically allow for a reliable decision. Document even an inconclusive test.
The test question must relate to the content
First, formulate a verifiable hypothesis. ‘A comparison by use case facilitates product selection and increases the proportion of unique product clicks’ is more precise than ‘We are testing a more attractive design’. This determines which element should be changed and which metric answers the question. A test of the entire brand requires a different framework to a comparison of two entry points.
A model example: A rucksack shop displays three models by price in variant A. Variant B organises the same models by day hikes, commuting and travel. Products, prices, delivery times and recipient rules remain the same. The question therefore concerns the selection guide. If a discount is added at the same time, it is no longer clear whether the new categorisation or the offer has caused the difference.
Set the target metric before launch
Choose the metric that answers your hypothesis: unique clicks for selection guidance, relevant orders for purchase impact, or completed sign-ups for a form question. Privacy features can affect opens, so they do not automatically prove attention. Set additional limits for warning signs such as unusual unsubscribes or complaints.
Klaviyo explains in the Guide to statistical significance how campaign tests are evaluated. In the specific test, check the winning metric used and the stated certainty. The economic significance is another matter: a small relative advantage may make little practical difference. You should therefore evaluate absolute values, denominators and effort together, rather than simply adopting the variant highlighted in colour.
Ensure variants are comparable and technically equivalent
Check both versions with the same level of care. Content, prices, landing pages and personalisation must work correctly. A broken link in Variant A turns the test into an error report rather than a fair comparison. Check on smartphones, desktops, in dark mode and for missing profile values. The respective change should appear in the received message exactly as intended.
Name the variants clearly and save the test plan. Note down the hypothesis, the section amended, the target audience, the expected outcome and the evaluation date. Do not alter current versions without notice. If a correction is necessary, document the change and decide whether the existing results are still comparable. A traceable restart may be preferable to a mixed test run.
Distinguish between ‘campaign test’ and ‘flow test’
A campaign reaches its target audience on a scheduled occasion. A flow test gathers observations as contacts enter the flow over time. The Klaviyo guide to testing a flow email describes the relevant functions. When doing so, check whether you are modifying only a single message or, in addition, entire paths and waiting times.
In the case of flows, the composition of the people involved may change over time. A sale, a new sign-up source or a modified checkout process will influence the data. You should therefore define a suitable observation period and document any significant parallel changes. Do not evaluate the decision until both variants have had comparable opportunities. A short-term click advantage does not necessarily mean that suitable purchases will follow later.
Avoid small sample sizes and premature termination
Two additional purchases can create a large percentage gap if the baseline figure is small. Therefore, place the absolute numbers alongside the rate. Do not continue evaluating until the desired variant takes the lead in the short term. Define in advance when the test is to be evaluated and how to deal with an ambiguous result.
Where the volume is low, a broader difference in content may be more informative than a minor colour change. Qualitative feedback can also aid understanding, but must be marked as such. A test without a clear conclusion is not a failure. It shows that the available data does not yet sufficiently support the change. In such cases, the simpler or more easily maintainable option may be a sensible editorial choice.
Interpreting the results using consistent measurement criteria
Compare the same metrics and complete time periods. Take into account attribution rules, automated interactions and varying order values. If the target metric was clicks, this remains the primary test question; a coincidentally higher single basket value must not subsequently become the new decision-making criterion without justification. Additional metrics help with context, but do not replace the plan.
Record the result, the uncertainty and the next action. For example: organising content by use case did not produce a reliable improvement in clicks, but it simplifies product advice, so it remains an editorial choice. This is more precise than claiming an unproven uplift. When a result is clear, check the implemented variant and its links again.
Checklist for the next A/B test
- A specific hypothesis describes the problem and the expected change.
- Variants differ at a clearly identifiable point.
- The target metric, denominator and evaluation time are fixed.
- Both the messages received and the landing pages have been checked.
- Parallel changes to the target audience, offer or tracking are documented.
- Absolute figures, uncertainty and economic significance are assessed.
Also retain tests without a clear winner. The documentation prevents the same team from treating a previous assumption as a new insight. Good experiments lead to transparent decisions and highlight the limitations of the available data.
Sources and further documentation
Product documentation and primary sources relating to the steps described. Editorial source date: 5 October 2026.