Machine learning marketing is the use of models that learn patterns from customer data to segment shoppers, predict behavior and personalize messages. Several vendor features have an eligibility floor: Klaviyo’s predictions require 500 customers with orders and 180 days of history, and Braze churn prediction typically needs 300,000 monthly active users.
This guide is written for ecommerce marketers deciding what machine learning can do for a store of their size, not for students of the field. It is published by Relvino, which sells an autonomous decisioning system, so read it as informed but interested. Every quotation below is from the vendor’s own page as served on October 5, 2026.
The definitions agree. Salesforce: “Machine learning is a branch of artificial intelligence that uses algorithms to imitate humans to improve accuracy in analyzing data, identifying patterns, and making predictions.” Braze: “Machine learning is a branch of artificial intelligence that trains systems to learn from data and improve over time”, without manual programming. Mailchimp’s guide says machine learning “consists of technology that receives input, identifies patterns, and independently adapts to new data in order to form solutions and solve problems.”
The use cases agree too. Salesforce lists “customer segmentation, predictive analytics (e.g., churn prediction), personalized recommendations, fraud detection, optimizing ad spend, and automating customer service.” Itransition’s list is customer segmentation and ad targeting, recommendation systems, marketing automation, content optimization, marketing analytics and predictive customer analytics. Aerospike gives an example: “One example is customer lifetime value (CLV) prediction. ML analyzes a customer’s past interactions and purchases to estimate the total value they will bring over their relationship with the company.”
All of those definitions describe what a model does once it has learned. None of the definition pages says how much a store has to show it first.
The data requirement lives in each vendor’s documentation, and it is specific.
Klaviyo predictive analytics. The help center lists the conditions: “At least 500 customers have placed an order.” It adds that “This does not refer to total profiles, but rather the number of people who have actually made an order with your business”, and that the orders must not be cancelled, refunded or zero-value. Also required: “You have at least 180 days of order history and have orders within the last 30 days”, and “You have at least some customers who have placed 3 or more orders.”
Google Analytics 4 predictive metrics. “In the last 28 days, over a seven-day period, at least 1,000 returning users must have triggered the relevant predictive condition (purchase or churn) and at least 1,000 returning users must not.” Meeting it once is not enough: “Model quality must be sustained over a period of time to be eligible”, and if quality falls below the threshold, “Analytics will stop updating the corresponding predictions and they may become unavailable in Analytics.”
Braze Predictive Churn. The troubleshooting page: “Typically, you need 300,000 Monthly Active Users in a single workspace.” It names two error states, “Not enough data to train” and “Not enough past non-churners to reliably build the Prediction”, and states the principle plainly: “Predictive Churn (and any machine learning model) is only as good as the data available to the model.”
Mailchimp. Its customer lifetime value and purchase likelihood features publish no numeric minimum: “You must have a connected online store” and “You must have sent at least one email campaign to your audience to generate predictive analytics.” One of its prediction types answers the floor a different way, which the section on whose crowd a model learns from comes back to.
The floor is not vendor caution. A prediction about a person is an estimate from people who looked like that person, and an estimate from few examples is noise. Google’s machine learning glossary makes the related point about overfitting: “Training on a large and diverse training set can also reduce overfitting.” Klaviyo says it about its own output: “predictions work best when averaged over many customers and are not expected to be exact for any single individual.”
That second sentence matters more for marketing than any threshold. A churn score or an expected next order date is reliable for a group and rough for a person, and marketing messages go to persons. The floors above are the population a vendor needs to train one model well. Marketing decisions are narrower than one model: which offer, on which channel, at which hour, for a shopper with three page views and one past order. Each of those decisions is a slice of the population, and a slice of a 500-purchaser store is small. That last step is this guide’s reading, not a vendor’s; the vendors publish the floor for the model, not for the decision.
Klaviyo says what its expected-date-of-next-order model does when one customer has too little history, and the answer is to borrow from other customers: “If the customer’s orders don’t exhibit a pattern or if we don’t have enough data on the customer, Klaviyo will make a reasonable prediction based on how your other customers behave.” And for the one-order case: “For one-time purchasers, since we don’t know much about their purchase behavior, we calculate their expected date of next order using data across all of your customers.”
That is sensible, and it shows where the information for a thin profile comes from: the crowd. Within one store, the crowd is that store’s customer list. For many of a store’s customers (first-time and one-order buyers) the per-person data is thin by construction, so what a model says about them comes largely from other customers. Profiles with no order at all get no such prediction from Klaviyo.
Each floor above is counted in the brand’s own numbers, and each of those features learns from one brand’s data inside one workspace or account. Mailchimp shows the other option for one prediction type: “Predicted demographics is based on shared Mailchimp account information from around the world, and uses the same gender and age range categories as Google Ads.” Relvino takes that route for its decisions: its Large Retail Model is trained on 7M+ data points across 10K+ retailers and is then adapted per merchant, so a new store’s first decisions come from that cross-retailer model rather than from its own history. Whether a pooled model fits a given store is an empirical question, answered by the 14-day pilot described in the Klaviyo vs Relvino comparison, not by the size of the pool.
Most of the use cases on the definition pages end in a number: a churn score, a predicted lifetime value, a segment. The number then waits for a person or a rule to act on it, which is the gap the next best action guide describes and the reason RFM scores and behavioral segments still need someone to decide what each one receives. A system that decides as well as predicts takes the score as an input and outputs the message, the channel and the time, or a decision not to send; Relvino makes that decision per shopper in about 80ms. The what is AI marketing guide covers the wider vocabulary, and send time optimization is a worked example of one model, built on per-person history, applied to one decision.
It is software that learns from past customer behavior to produce an output a marketer would otherwise guess: a segment, a churn score, a predicted next order date, a product recommendation, or a choice of message. The test of a feature is what it outputs: a score, a segment or a predicted date, or the message itself.
Each vendor publishes its own floor in its help center, and they are counted in different units and differ widely: hundreds of purchasing customers for Klaviyo’s predictions, a thousand returning users on each side of the outcome for GA4, and typically hundreds of thousands of monthly active users for Braze’s churn model. A store can check its own numbers against those pages in a few minutes, and that check says more than any demo.
Both. ChatGPT is built on a large language model, and large language models are trained with machine learning, which is a branch of artificial intelligence. In marketing the useful split is not AI versus ML but what a model is for: generative models write copy, predictive models estimate what a shopper will do, and a decisioning system uses those estimates to choose what to send, if anything.
Our view, as a company that sells autonomous decisioning: the per-customer execution work (building segments, maintaining flows, scheduling sends) is the part software is taking over. The parts that stay human are the ones a model cannot learn from data: the offers a brand is willing to make, its margin floors and quiet hours, its voice, and the judgment about what to test next.
Count customers with at least one non-cancelled, non-refunded order and note the date of the first order. Under 500 such customers or 180 days of history, Klaviyo will not show its predictive section; for GA4, check how many returning users purchased, and did not, in the last 28 days. Then list the decisions the store wants a model to make and ask each vendor what data that specific decision needs.
About 30 minutes for the technical cutover: connect the Shopify store, point the sending domain, connect SMS, and shopper data ingests automatically. Revenue is then proven in a 14-day pilot run beside the existing setup and scored on revenue per shopper.