Incrementality testing in marketing measures how many conversions a message caused, by comparing a randomly chosen group that received it with a holdout group that did not. In lifecycle email and SMS it matters because the default attribution rules at Klaviyo, Omnisend and Attentive credit orders after an automatic Apple Mail open or a text that was only delivered.
This guide covers incrementality testing for the lifecycle program: the email and SMS flows and campaigns a store sends, and the “attributed revenue” figure reported for them. It quotes three vendors’ help centers as served on October 8, 2026, then explains the test, the lift formula and what to do with the answer. It is published by Relvino, which decides per shopper across email, SMS and pop-ups, so read it as informed but interested.
A conversion is incremental when it would not have happened without the marketing. The opposite is a conversion the marketing was merely near: the shopper was coming back anyway, and a message happened to land first. Attribution reports cannot tell those two apart, because they assign credit by rules about timing. Incrementality testing separates them by withholding the marketing from a random group and counting what that group does on its own.
Klaviyo describes its own version of the method in those terms. Its holdout feature, in Klaviyo’s words, provides “a way to exclude a number of profiles from your messaging to determine the incremental success of your marketing program.” The platform that reports attributed revenue on every flow ships a separate tool for the question attributed revenue does not answer.
All three platforms publish their attribution rules in their help centers. Read side by side, the rules share three features: credit goes to the last message touched, a message is eligible for days afterwards, and the touch that qualifies does not have to be a person reading it.
Klaviyo’s message attribution article states the default: “By default, Klaviyo uses a last touch attribution model with the following lookback window for all new accounts”. The list that follows begins “5 days for email clicks”, “5 days for email opens” and “5 days for text messages clicks”, and includes “12 hours for text message deliveries”. Its older conversion tracking article summarizes the same settings more loosely (“the default window is 5 days for email and SMS and 24 hours for push”), so the finer list is the one to check against an account. The two also disagree on texts: the older article says conversions do not count when someone “just received the message but made no further action.”
On opens, the conversion tracking article is explicit: “Conversions are tracked if someone opens or clicks, or if Apple Mail Privacy Protection (MPP) automatically opens the email.” Excluding those opens is a setting a brand can change. What an MPP open is, Klaviyo explains in a separate article: “We have no way of distinguishing between a true human open and an automated open when MPP is enabled.”
Omnisend’s defaults are “Email: 7 days” and “SMS: 24 hours”, with the same last-touch rule. Its FAQ answers the open question directly: “Apple MPP opens are included as engagement for sales attribution by default for all accounts.” Two further rules widen the net. In-store orders count (“If a customer engages with your email and makes a purchase in-store within the attribution window, that sale will be attributed to the campaign”), and refunds do not come out (“Refunds and cancellations do not impact attributed sales metrics”).
Attentive lists four kinds of engagement it credits, and the second needs no action from the shopper: “Delivered text messages: A subscriber receives a text message, doesn’t click a link in the message, and then makes a purchase.” For new customers, its default delivered window for texts is 24 hours, rolling. A unique coupon code can also earn credit “even if the purchase occurs outside of other defined attribution windows.” (Attentive’s help-center page refused automated requests on October 8; the text above is the same article served by its help-center API.)
Put those rules together and attributed revenue becomes a precise description of timing: an order placed within a few days of a qualifying touch. That is a different quantity from caused revenue, and the gap between them has a predictable shape.
Attributed revenue still ranks campaigns consistently under one set of rules. It does not say whether a flow earns its sends. Timing problems compound it: a send scheduled for when a shopper reads mail (the subject of the send time optimization guide) collects credit on any order that follows, read or not.
In Klaviyo the windows live under Settings, then Attribution, where each channel expands to show its touchpoints. The help center notes that “If you do not select a channel checkbox, Klaviyo will exclude this particular interaction from counting towards attribution”, and that a Compare model tool previews how a change shifts attributed revenue before it is applied. Before any test, write down the current settings, because the attributed figure the test is compared against depends on them.
The method has three parts. First, a random share of eligible shoppers is assigned to a holdout (also called a control group) that does not receive the message, the flow or the program being tested. Random assignment is the whole point: if the holdout were chosen by hand, or were simply the shoppers who did not open, the two groups would differ before the test began. Second, both groups run through the same period, so seasonality and promotions hit them equally. Third, the comparison uses every order each group placed, not attributed orders. Klaviyo’s holdout report makes the same distinction in its own words: “Note that is total conversions made by your profiles and not attributable conversions.”
Lift is the relative difference in conversion rate between the two groups:
Lift = (treatment conversion rate − control conversion rate) ÷ control conversion rate
A hypothetical abandoned checkout flow, with round numbers chosen for the arithmetic, not taken from any store:
Now suppose the platform’s report for the same flow and period showed 700 attributed orders. The test says the flow caused about 180 orders in total, so at least 520 of the 700 credited orders would have happened anyway. The flow works, and the attributed figure overstates it by close to four to one. That ratio, incremental over attributed orders, is the number worth keeping from a test: it turns every future attributed report for the flow into an estimate of what it caused. Abandoned checkout is a natural first test, since those shoppers were mid-purchase (see abandoned cart vs abandoned checkout).
A one-point gap is easy to produce by chance in small groups, so the result needs a confidence measure. Klaviyo reports a “Win probability” for its holdouts, recommends “about 5% of your total marketable profiles” as the holdout size and recommends running for 3 months. It also warns that “ending a test early may result in insufficient or skewed data.”
The two are often confused, and Klaviyo’s holdout article separates them cleanly: holdout groups measure “the impact of your marketing communication by not sending anything to a group of people. In contrast, A/B testing allows you to compare 2 or more versions of your marketing communication.” An A/B test answers which version of a message wins. It cannot answer whether either version beats sending nothing, because both arms receive a message. A subject line test can crown a winner in a flow whose true lift is zero.
Klaviyo’s built-in tool is a global holdout: one at a time, across every channel (“it cannot be turned on for one specific channel vs. another”), and available only to accounts with at least 400,000 profiles in total. Flows a brand chooses to send to the holdout anyway, such as a welcome series, reduce what the test can see; Klaviyo notes “You may see a decrease in overall marketing uplift as the holdout group is still receiving some messaging.”
A smaller program can still run the method by hand, flow by flow. The usual approach is to assign each profile a random value once (a random number stored as a profile property works on most platforms), exclude a fixed slice of those values from one flow, and compare total orders for the two slices over the same weeks. That is a method, not a documented vendor feature, so it is worth checking three things before trusting the result:
By the same logic, the gap should be widest in flows whose audience already buys often; browse flows are a candidate, because their audience is shoppers the store already knows (see the browse abandonment email guide).
Relvino decides per shopper, in ~80ms, whether to send anything, and if so what and where, across email, SMS and pop-ups, inside guardrails the brand sets. A decision not to message a shopper is part of the design rather than a gap in a flow, which is the behavior a holdout test rewards. How that compares with a flow-based Klaviyo setup, and the 14-day pilot used to measure it, is in the Klaviyo vs Relvino guide.
A finished test lands in one of four places, and each points to a different action:
Whatever the outcome, record the attribution settings in force during the test beside the result. The ratio only holds while those settings do, and the platform will recalculate the attributed half of it the next time someone changes them.
Incrementality is the share of an outcome that happened because of the marketing and would not have happened without it. It is always measured against a counterfactual, usually a randomly selected group that did not receive the marketing. The term applies to any outcome, not only orders: incremental sign-ups, incremental repeat purchases or incremental revenue per shopper are all measured the same way.
In marketing, incremental testing is another name for incrementality testing: a controlled experiment that withholds a message, flow or channel from a random group and compares outcomes. (In software, the same phrase means testing modules one at a time as they are integrated, which is unrelated.) Marketing teams also call the variants holdout tests, lift tests or, when whole regions are held out instead of individual people, geo tests.
An A/B test splits recipients between two versions of a message, so every participant is marketed to. An incrementality test splits them between a message and no message. The two can run together: hold out a slice of the flow’s audience, then A/B test versions inside the remaining treatment group. The winning version then has a known lift over silence, not just over the losing version.
Divide the difference between the treatment and control conversion rates by the control rate to get lift as a percentage. Multiply the difference in rates by the number of treated shoppers to get incremental conversions. For revenue, run the same arithmetic on revenue per shopper in each group rather than on attributed revenue, since the control group has no attributed revenue by design.
Its conversion tracking article says conversions are tracked when Apple Mail Privacy Protection automatically opens an email, and its attribution settings page lists MPP opens among the interactions a brand can exclude. Klaviyo adds a caveat: excluding MPP opens removes them from attribution but does not include removal from reporting, so open rates in reports still include them.
About 30 minutes for the technical cutover: connect the Shopify store, point the sending domain, connect SMS, and shopper data ingests automatically. Revenue is then measured over a 14-day pilot. Keep a record of the old platform’s attribution settings before switching, so attributed figures from before and after the move are compared under the same rules.