Three Simple Rules to Drive Results with AI
What We Learned Building Performance Marketing Agents for eCommerce
Mark McMaster is speaking on the keynote panel "Proof Not Promises: Real-World AI Wins" at eTail Boston on August 11. This blog captures what we learned over the past year building AI for eCommerce and retail applications.
There's no shortage of buzz in the retail and eCommerce space around AI: it can drive personalization, make marketing more efficient, streamline customer service, better leverage data, and forecast demand. All of it sounds worthwhile. But how do you know where to prioritize AI, and how do you know it will deliver results?
This is a question we've had to answer directly while building performance marketing agents like MAI to manage and scale campaigns across Google, Meta, and other marketing platforms. Performance marketing is a data-intensive function which we know is an ideal use case for AI, but which decisions should agents actually be trusted to make, and which ones still need a human? Here's what we've learned.
1. Apply AI where complexity drains resources
We've found that the situations that benefit most are the ones where complexity is already eating time and budget faster than a team can manually keep up. In managing search and social campaigns, a few patterns show up consistently.
High-SKU catalog accounts are one. The sheer combinations of allocating budget across hundreds or thousands of products outpaces what a person can manually track and adjust in real time. Our work with retailer Nutrition Faktory showed this clearly: while past efforts struggled to deliver profitable sales across their whole product assortment, MAI now optimizes across thousands of SKUs to their actual margin profiles, not platform ROAS, and the account now runs at 3x profitable ad spend with over 4x ROAS.
Search intent is a second pattern, and it compounds the first. Customers discover the same product in thousands of different ways, and manually mapping keywords to intent doesn't scale. Velotric sells e-bikes, a purchase people research heavily before buying. MAI's keyword harvesting workflows for Velotric ran persistent exploration and found relevant new traffic the team hadn't tapped: cost down 9%, conversion value up 146%, ROAS up 172%. Meanwhile, working with an AI-enabled privacy app, we found the opposite problem: a category so new that search language was still forming. Capturing that non-brand intent early drove a significant increase in search purchases and cut CPA by half.
Multi-channel accounts are another. Running Google, Meta, and Microsoft together, the actual bottleneck isn't managing any single platform well, it's reconciling signals across all of them at once. We worked with a DTC home brand to transition from monthly media mix reports, which reallocated budget based on a historical snapshot, to real-time cross-channel optimization built on multi-touch attribution. Both ROAS and total revenue went up.
2. Put first-party data to use in agentic workflows
The next pattern is less about the account structure and more about what's feeding it. Most brands are sitting on first-party data they aren't putting to work. Ask any in-house marketing team about their biggest frustration with reporting and you'll hear some version of the same thing: every platform reports its own numbers, and none of them agree. Google says one thing, Meta says another, and reconciling the two eats hours every week with no real source of truth at the end of it.
First-party data helps fix the attribution problem. A brand can look at its own numbers instead of trusting what each platform says about itself: real profit margins from Shopify, not just revenue, plus in-site analytics, smaller signals like site search or time on site that hint at intent early. It can also see the same customer across channels instead of double-counting them as separate results. Most brands already have this data sitting there. They're just not using it yet.
Our work with a yoga brand shows what that looks like in practice. Using the brand's own data, MAI agents ran incrementality measurement on branded search, validated what that traffic was actually worth at the margin, and set the brand search budget that maximized ROAS across Google. That's a question platform-reported numbers can't answer, because every platform grades its own homework.
The same data unlocks full-funnel measurement. For high-consideration products, such as e-bikes and wedding rings, our customer journey monitoring showed which touchpoints actually start the path to purchase, not which one happened to collect last-click credit, and validated it with incrementality testing instead of platform attribution.
3. Marketing is both art and science. Let humans curate the art
This is probably the most telling decision we've made at MAI, precisely because it goes against the instinct to automate everything possible. Creative doesn't have what the other decisions above have: a fast, unambiguous feedback signal. Whether a piece of creative is good isn't something a model can score the way it scores a bid adjustment. It's a judgment about taste, brand fit, and tone, and those don't reduce cleanly to a number a system can optimize against. It's the art side of the equation, not the science side, and knowing the difference matters.
So we didn't automate creative generation. Not because the technology can't generate creative, it can, but because generating creative and being trusted to make creative decisions unsupervised are two different bars. MAI separates what a model is capable of from what it should be trusted to decide unsupervised, and creative sits firmly on the "assist, not decide" side of that line. That's a deliberate choice to protect the human element of a brand, the part AI shouldn't be the one to define.
That doesn't mean AI sits out of creative entirely. There's a science side to creative too, and that's where agents earn their keep: grouping assets by performance so winners surface faster, carrying winning concepts from Meta into other platforms' asset groups and back, and detecting creative fatigue so assets rotate before performance decays. A fast-fashion brand we work with used these insights to help new creatives scale more quickly on Meta, covering more new products and accelerating sales of new SKUs. Humans made the creative. The system decided how fast and how far each piece traveled.
The actual dividing line
Put these threads together and a pattern shows up. The decisions worth automating are the ones with fast, consistent feedback loops and low individual stakes. The business processes worth prioritizing are the ones where complexity, whether in SKU count, channel count, or the first-party data most brands are sitting on and not using, is already draining resources faster than a team can keep up manually. And the decisions worth keeping human are the ones where a mistake is expensive and slow to correct, or where "correct" is itself a matter of judgment rather than measurement.
From our experience, the brands earning genuine ROI from AI aren't the ones who handed everything over. They're the ones who knew exactly where to draw the line, and built a system built to respect it.