AI marketing agents are software systems that take a marketing goal and carry it out: drafting, segmenting, deciding and sending. The useful way to sort them is not by model or feature list but by blast radius, meaning what each one is allowed to do to a customer before a person sees it. Four tiers cover the whole market.
We build the fourth tier, autonomous decisioning for ecommerce, so read this as informed but interested. The first three tiers are described fairly, and most marketing teams should be using one of them this week. The sorting rule is ours.
Salesforce defines AI marketing agents as “specialized software systems that autonomously reason through data, make decisions, and execute marketing tasks like segmentation, personalization, and campaign activation.” That is a good definition, and most products that carry the label do not meet it. Most products called agents draft, suggest or summarize, and a person still decides what reaches a customer. Nothing is wrong with that. It is simply a different amount of autonomy, and the amount is what a buyer needs to know.
The gap between the definition and the products is why feature comparisons of AI marketing agents are so hard to use. Two tools can both say they generate campaigns, predict the best send time and personalize content, while one of them never sends a message without a click and the other sends thousands a day that nobody reviewed. The question that separates them is blast radius: what is this agent permitted to do to a real customer without asking?
ChatGPT, Claude and Gemini, plus the assistants built into commerce tools such as Shopify’s Magic and Sidekick. They draft subject lines, product descriptions, briefs and reports at a person’s direction. Nothing reaches a customer that a person did not paste, schedule or send. They are the safest kind of AI in marketing, and they change the cost of writing, not the quality of the decision about who gets what.
Klaviyo ships K:AI, whose Composer will, in Klaviyo’s words, “build the campaign for you”; ActiveCampaign’s Active Intelligence, in its own words, “spots upcoming moments and builds an entire campaign for you” and delivers “Campaigns drafted. Issues caught. Insights delivered.” Salesforce’s Agentforce for Marketing describes the split as marketers who “Set the strategy, define the guardrails, and hand off execution.” In their own words these products draft campaigns, build audiences and adapt journeys. Where they sit on the blast-radius scale is our judgment: the co-pilot line is the point at which a person still designs the flow the agent works in and reviews what it drafts. Its blast radius is bounded by that journey: it can make a designed flow better, and it cannot decide that the flow was the wrong idea for this shopper.
Bidding and budget agents on Meta and Google, support agents that answer tickets, agents that reorder inventory. These do act without approval, which is what makes them real agents, and their blast radius is deliberately narrow: a spend cap, a channel, a category of ticket. A bidding agent can waste a day’s budget. It cannot email your best customer three times in an afternoon.
The fourth tier decides, for each shopper and in real time, whether a message is warranted at all, and if so which offer, on which channel, at what moment, and then sends it. Nobody wrote a rule for that shopper. This is the only tier where the definition above is met in full, and it is the only tier where the trust question is serious, because the thing at risk is not a budget line but a customer’s patience. That is why the buying criteria for tier 4 are guardrails and observability, not model quality. We covered how to audit such a decision in AI marketing for Shopify.
Three questions sort any product called an AI marketing agent, and none of them appears on a feature grid.
Sorted this way, the market looks less crowded than the listicles suggest. Most named products are tier 1 or tier 2, which is why a team can adopt them in an afternoon and why they leave the plateau untouched: the flow paradigm that decides who gets what is still hand-built, and the agent works inside it. The category itself is defined in agentic marketing, and the buying checklist for tier 4 is in agentic marketing platform.
Relvino runs the loop Observe → Decide → Act on a store’s owned channels. It watches live shopper signals, decides offer, timing and channel for that one shopper in about 80 milliseconds, and executes across email, SMS and on-site. Guardrails are set once: margin floors, channels, brand voice. Inside them, 100% of flows run without a human in the loop, which is the product claim and the whole point.
The priors come from a Large Retail Model trained on 7M+ data points across 10K+ retailers and 1.78M shoppers, so the agent knows what works in a category before it knows a store. Brands replacing an incumbent see 2–6× ROI in 30 days and up to 10× revenue uplift year over year; Stein Mart saw 6X ROI in first 14 days. Because the agent decides whether to send at all, the registered efficiency claim is 80% less spam and lower send costs. One cleared example: POV Beauty, 2X fewer emails, same revenue. The observability surface is a live view of the interventions as they run, which is the most-used screen in the product, and it exists because a tier 4 agent has to be auditable to be trusted.
If the problem is writing, buy tier 1 today. If the flows are built and the problem is keeping them tuned, tier 2 is already inside the platform you pay for; turn it on. If the problem is paid media, tier 3 is mature and the vendors’ own agents are usually the best ones. If the problem is that lifecycle revenue has been flat for three quarters while the team maintains dozens of flows, none of the first three tiers touches it, because they all leave the decision about who gets what to a person in a builder. That is the tier 4 problem, and the right test for a tier 4 agent is not a demo but a bounded pilot: run it beside the incumbent for 14 days and compare revenue. Autonomous email marketing walks through what that looks like on the email channel.
Yes, and the honest answer depends on which tier. Assistants such as ChatGPT draft marketing but send nothing. Co-pilots inside platforms such as Klaviyo or ActiveCampaign draft and build campaigns for a person to review. Task agents bid and answer tickets inside a cap. Autonomous decisioning systems such as Relvino decide and send per shopper inside guardrails, which is the only tier that runs marketing rather than assisting it.
There is no single best, because the tiers solve different problems. For writing, a general assistant. For tuning existing flows, the co-pilot already inside the platform. For paid media, the ad platforms’ own agents. For lifecycle marketing that has plateaued under hand-built flows, an autonomous decisioning system, judged by its guardrails, its observability and a side-by-side pilot rather than a demo.
It follows the tier. Assistants are priced per seat. Platform co-pilots are bundled into contact-based plans and often gated by tier; ActiveCampaign, for example, lists Active Intelligence as limited access on its lower plans and unlimited on Pro and Enterprise. Task agents are usually included with the ad or support platform. Autonomous decisioning is priced differently; Relvino runs 3–7× cheaper than Klaviyo, with details on its pricing page.
Asked generally, people usually mean the general-purpose assistants: ChatGPT, Claude and Gemini. For marketing those are tier 1, which is to say they draft and analyze but do not act on customers. The top agents for a marketing team are better named by tier than by brand: the assistant it writes with, the co-pilot inside its platform, and, if it wants decisions made rather than drafted, an autonomous decisioning system on its owned channels.
The technical cutover takes about 30 minutes: authenticate the sending domain, install the Shopify app, and historical data ingests on its own. Nothing is rebuilt, because there are no flows to recreate. Revenue runs on a different clock: run a 14-day pilot with Klaviyo paused rather than cancelled, and compare the two side by side before deciding.