Est.

Selection Bias in Engagement-Based Lift Claims

Selection bias makes high-intent shoppers appear more valuable to AI than they actually are.

Reporter · · 8 min read
Cover illustration for “Selection Bias in Engagement-Based Lift Claims”
Incrementality Testing · September 20, 2026 · 8 min read · 1,900 words

DTC brands keep repeating a number that supposedly proves AI chat is quadrupling conversion: 12.3% for shoppers who engage with an AI widget versus 3.1% for those who don't, a gap first published in an AI ecommerce shopper behavior report and since republished across an industry roundup and a stack of operator guides. The math is right. The conclusion that AI causes that gap is not. What the stat actually captures is who chooses to open a chat window in the first place, and that selection problem quietly rewrites the whole story.

How self-selection produces the gap without AI's causal contribution

Start with the question the stat skips entirely: who initiates a chat conversation to begin with?

Shoppers who click into a chat widget aren't a random sample of site traffic. They tend to be further along already, comparing two sizes, checking a return policy, deciding between the last item in stock and a similar one. That's higher intent walking in the door before a single word gets typed. The AI is responding to a signal that already existed in the shopper's head. Calling the resulting conversion lift "AI-driven" mistakes the messenger for the cause.

There's a useful parallel in paid media attribution. Two ad platforms will each claim full credit for the same sale when a shopper sees an ad on one platform, then later clicks an ad on another platform and buys. Both platforms report 100% attribution, and add them together and you get credit for well more than the one actual sale that occurred. Comparing engaged versus non-engaged shopper cohorts runs the same trick in a different outfit: it hands AI full credit for intent it didn't create.

The trap gets expensive when a support bot that pushes up time on site while quietly lowering completed checkouts looks like a win on an engagement dashboard and loses money in the P&L. A support bot that pushes up time on site while quietly lowering completed checkouts looks like a win on an engagement dashboard and loses money in the P&L. Engagement is not revenue. Test AI against conversion rate and revenue per session, not against how long someone lingered on the widget or how many messages they sent.

Forrester's 2025 research on what it calls the "personalization paradox" gets at the mechanism directly: automation amplifies whatever signal already exists in the input. If a chat tool activates more often on high-intent sessions, which most tools do by design, then any lift metric built on engaged-versus-non-engaged comparison is structurally inflated before a single customer ever sees a recommendation. No further numbers are needed to make that case. The logic carries the argument on its own.

The same bias in AI referral traffic data

The identical bias appears one layer downstream, in traffic that arrives at a store from an AI referral source instead of a chat widget on-site.

Adobe's Q1 2026 data found that AI-referred visitors convert 42% better than non-AI traffic, spend 48% longer on-site, and generate 37% more revenue per visit. Those figures are accurate. They're also describing a pre-sold customer. By the time someone clicks through from a conversation in ChatGPT or Perplexity, they've usually already done the comparison shopping inside that chat. They arrived with a decision half made and a specific product in mind.

AI referral traffic is the filtered remainder: people who knew roughly what they wanted, asked a model, got pointed somewhere, and clicked. It's the filtered remainder: people who knew roughly what they wanted, asked a model, got pointed somewhere, and clicked. Everyone still undecided got filtered out before they ever touched the site.

The site-level numbers show the fingerprints of that filtering. IRP Commerce platform data put June 2026 conversion at 2.03%, up from 1.85% a year earlier, but visitor volume over that same year dropped 12.23% while revenue per session climbed 28.48%. Read together, those three numbers show the conversion rate is rising because low-intent browsers stopped showing up, not because the site got better at converting browsers. It's rising because low-intent browsers stopped showing up. The denominator shrank. The numerator didn't do all the work the headline conversion rate implies.

That distinction matters for anyone crediting an on-site AI tool with a conversion gain that's really an audience effect happening upstream, inside a model's own recommendation logic, long before a shopper lands on the store.

Diagram: The Shrinking Denominator: Why Conversion Rates Rise Without AI Doing the Work. Visualizes: Visualize a before/after split showing how the IRP Commerce June 2026 platform data tells three simultaneous stories: visitor volume fell 12.23%…

Why vendors keep reporting the stat this way

None of this means vendors are cooking the books. It means the incentive runs one direction: report the comparison that looks best, and engaged-versus-non-engaged is the easiest split to pull from a vendor's own dataset, no client cooperation required.

A genuine randomized holdout is a hard sell. It means deliberately showing some high-intent shoppers a worse experience on purpose, just to measure the difference, and no vendor wants to walk into a renewal conversation proposing that to a client. So the comparison that gets built and shipped is the one that's easy to run and flattering to report.

Paid media already lives with a version of this problem at scale. Meta's default attribution window credits a sale within 7 days of a click, 1 day of an engage-through, or 1 day of a view-through. Google runs an equivalent model. Every platform separately claims its own total, and the sum of those totals routinely exceeds what the store actually made in a given period. AI commerce tools inherited that same reporting culture wholesale, just with a chat window instead of an ad unit.

The industry knows the gap exists. A survey from EMARKETER and TransUnion found 52% of US brand and agency marketers now run incrementality testing, a sign of how far standard attribution has already fallen short even in the more mature world of paid media measurement. AI-driven commerce, still newer and less standardized, has further still to go.

What a credible measurement baseline requires

Real incrementality testing means geo holdouts or A/B structures that isolate what the AI tool actually contributed, the only reliable way to answer whether those conversions would have happened anyway.

Short of a full holdout, a more honest version of the engaged-shopper comparison measures AI-engaged shoppers against the brand's own historical baseline for similarly high-intent sessions, not against the entire pool of non-engaged visitors who never showed comparable intent. That's an apples-to-apples comparison instead of a mismatched one.

Credible personalization lift estimates tend to be far more conservative than the multiples-of-lift claims that circulate in vendor content and roundup articles, and a realistic target is a far more honest one for a brand actually deploying AI to hit.

The right things to measure are conversion rate per session and revenue per session, tracked for cohorts the AI actually touched against matched cohorts it didn't. And even incrementality testing isn't a magic fix on its own: a report from Skai and the Path to Purchase Institute found 44% of marketers already question the reliability of the incrementality results they're getting. Running the test matters, but so does transparent methodology behind it. A blunt but useful gut check for any operator asks whether a vendor can show the conversion rate for sessions where the AI didn't engage at all; without that, the lift claim can't be falsified, and a claim that can't be falsified isn't worth building a budget around.

Why this measurement problem will get harder as agentic commerce scales

The measurement gap is about to matter a lot more, because the volume behind it is about to grow fast.

The Braze Retail Customer Engagement Review projects consumer adoption of agentic shopping jumping from 19% to 46% by the end of 2026. An IBM Institute for Business Value study found 45% of consumers already use AI for some part of the buying journey today. Adobe Analytics clocked a 393% year-over-year rise in AI-referred traffic to US retail sites in Q1 2026, a number that's actually down from 673% growth back in December 2025, but still a volume large enough to make sloppy attribution a real cost, not a rounding error.

That growth contains a distinction that most lift claims gloss over. Chatbots answer questions. Agents act: they read context, make decisions, and complete purchases without handing the shopper back to a search results page. A lift number measured against a legacy chatbot deployment doesn't transfer cleanly to an agent that's completing checkout on a shopper's behalf. The mechanism producing the number has changed.

OpenAI and Stripe's ACP and Google and Shopify's UCP are explicitly designed to be complementary infrastructure, built to work alongside each other rather than as mutually exclusive alternatives. ChatGPT Shopping research is now live for all logged-in US users, though Instant Checkout was discontinued in March 2026. Etsy and over a million Shopify merchants are already live on this infrastructure. Discovery is happening in places that don't show up cleanly, or sometimes at all, in a standard analytics dashboard.

That shift changes what brands are actually competing on. Selection into an AI agent's shortlist runs on data quality and catalog completeness, not on design polish or ad spend. A brand that reads the old multiple touted for AI-driven conversion, concludes its AI strategy is already working, and skips investment in catalog quality is making the wrong bet at exactly the wrong moment to make it.

What DTC brands should build toward instead

It's whether AI moved shoppers who wouldn't have converted anyway. It's whether AI moved shoppers who wouldn't have converted anyway, and that's a genuinely different question with a genuinely different answer.

Structured, enriched product data determines whether a brand gets cited and chosen by AI shopping surfaces, and it holds its value no matter which protocol or vendor ends up winning the standards fight over the next few years. Catalog quality compounds. A clever prompt or a slick chat interface doesn't, not in the same way.

A single AI system trained on a brand's own catalog, policies, reviews, and voice, deployed consistently across the brand's own site and the external AI surfaces where shoppers now go looking, produces measurement that's at least internally consistent. A patchwork of disconnected point tools, each reporting its own flattering slice of the funnel, fragments attribution further and makes the real number harder to find.

Speed to deployment is a real, practical criterion too, not a vanity metric. IDC's 2025 benchmarks put the median enterprise deflection deployment at 4.5 months. A brand agent that's live the same day, without a theme rebuild or a heavy IT lift, gives an operator a known starting point to run an incrementality test from immediately, rather than waiting months for implementation before the first real measurement is even possible.

The honest platforms are the ones willing to report AI-driven impact against a brand's own historical baseline for comparable high-intent cohorts, not against the entire non-engaged population. That's a harder number to report and a much more useful one to act on.

None of this argues for waiting around for perfect methodology before doing anything. With agentic shopping adoption headed toward 46% by the end of 2026 and AI referral traffic still growing at triple-digit rates, brands that wait risk losing share to competitors already capturing that traffic. The fix is building the measurement discipline that tells a brand, in its own data, what its AI actually did. It's building the measurement discipline that tells a brand, in its own data, what its AI actually did.

Sources

  1. Agentic commerce - Wikipedia
  2. paz.ai
  3. digitalapplied.com
  4. braze.com

More in Incrementality Testing