Est.

Ghost Ad Holdout Design for DTC Paid Social

Meta's ghost ads isolate true campaign impact when attribution claims prove inflated by two-thirds.

Features Editor · · 15 min read
Cover illustration for “Ghost Ad Holdout Design for DTC Paid Social”
Incrementality Testing · September 22, 2026 · 15 min read · 3,289 words

Every platform selling ad space has a financial reason to tell you the ad worked. That's not a scandal; it's just the incentive structure, and it means the number on a Meta or TikTok dashboard is not a measurement of causation but a claim of credit. A landmark Meta internal study across 15 advertisers found that only 34% of conversions the platform attributed to its ads were actually incremental, so two-thirds of the "results" a media buyer might report up the chain would have happened anyway. This piece walks through the fix: how holdout testing actually isolates causal lift, which of the three real designs fits a given budget and question, and where teams predictably get the sizing, duration, and timing wrong.

Consider the classic retargeting story. A high-intent shopper visits a product page, leaves without buying, sees a retargeting ad two days later, and completes the purchase. The platform logs that sale as a campaign win. But the shopper had already added the item to a mental cart; the ad may have done nothing but remind someone who was going to come back anyway. Multiplying that scenario across a catalog and a media budget makes the gap between attributed performance and actual causal impact the single biggest distortion in DTC media planning.

The structural reason for this over-reporting isn't a bug in any one platform's tracking; it's built into how each one is scored. Meta leans on view-through conversions and modeled data to fill gaps left by signal loss. Google, Meta, and email platforms all independently claim credit for the same transaction, because each system is built to justify its own budget line, not to divide credit fairly with its neighbors. Google shut down most Privacy Sandbox features in October 2025, leaving third-party cookies as Chrome's de facto default rather than replacing them with anything more private, and a mobile platform's app-tracking permission requirement and the collapse of third-party cookie alternatives turned signal loss from a temporary inconvenience into the permanent operating condition of digital advertising. There is no "after cookies" phase coming where clean attribution returns.

The financial stakes of getting this wrong have gone up, not down. Segwise.ai reports that average ecommerce ROAS fell to 2.87:1 in 2025 while Meta CPMs rose 15 to 22% across most verticals. When cost per impression climbs and reported returns compress at the same time, the gap between what a dashboard says and what's actually happening in revenue becomes a direct profitability problem, not an academic one. Solving it requires a methodology built to measure causation instead of correlation. That's what incrementality testing, and specifically the holdout test, exists to do.

What incrementality testing measures, and how holdout tests produce it

Incrementality testing isolates the causal effect of an ad by comparing two groups that are otherwise identical: one that saw the ad (treatment) and one that didn't (holdout, or control). The difference in outcomes between the two groups is the incremental lift. That's the entire mechanism. Attribution models never ask what would this person have done if the ad never existed. A holdout test answers exactly that, because the control group lived out that counterfactual in real time.

The arithmetic is simple once the groups are genuinely comparable: incremental revenue equals revenue from the test group minus revenue from the control group. The catch is in that qualifier, "genuinely comparable." If the two groups differ in some way that also affects purchase behavior, the subtraction produces a fiction dressed up as a data point.

The incrementality gap is where the practical damage happens. Say 5% of an audience was always going to buy, holdout or not. If the exposed group converts at 6%, the true incremental lift is just 1 percentage point, not the full 6% a naive attribution report would credit to the campaign. Every cost-per-acquisition figure built on that inflated 6% is wrong by a factor of six, and every budget decision built on that CPA compounds the error.

An A/B test compares creative variants within an audience that's already being advertised to, answering "which ad works better," not "does advertising work at all." A holdout test is a method that isolates causal lift; an A/B test compares creative variants within an audience that's already being advertised to, answering "which ad works better," not "does advertising work at all." It's also not the same as simply turning a campaign off and watching what happens, the so-called time-based on/off test. That design is the most commonly run and the least valid, because organic momentum, seasonality, and word-of-mouth don't pause just because the ad spend did. Whatever revenue happens during the "off" period gets misread as evidence about the ad, when it may just be evidence that customers kept buying regardless.

None of this confusion is limited to a handful of undisciplined teams. Roughly three in four marketers say their measurement systems lack the speed, accuracy, or trust they need to make confident decisions, according to the IAB and BWG Global's "State of Data 2026: The AI-Powered Measurement Transformation." Holdout testing is the controlled experiment that attribution, by construction, cannot be. But running one well is where most teams stumble, and the design choice is the first place that failure starts.

The three holdout designs DTC teams use on paid social: ghost ads, PSA holdouts, and geo experiments

Ghost ads work by having the platform withhold ad delivery from the control group while still "ghost-bidding" on their behalf in the auction, logging a phantom impression at the exact moment a real ad would have served. That auction symmetry is the whole point: it means the test doesn't waste spend on placeholder creative and doesn't distort how the algorithm learns, since the auction behaves as if every user were eligible. Meta's Conversion Lift product and The Trade Desk both offer this kind of randomized, auction-aware control, though it isn't available on every channel. The upside is speed and a low direct cost to run; the downside is that the test lives entirely inside one platform's black box, and there's no way to audit the holdout logic from outside.

PSA-based holdouts take a more visible approach. Users get randomly split into a group that sees the real ad and a group that sees a neutral public-service-style message instead, one designed to have zero effect on purchase intent. This preserves auction pressure while giving a genuine randomized controlled trial at the user level. The craft challenge is creative neutrality: a PSA that accidentally primes someone toward the category, even subtly, contaminates the entire control group. Industry practice generally puts PSA spend at 10 to 30% of the total test budget, with the exact split driven by how much statistical power the test needs. This design is workable across Meta, TikTok, Snapchat, and Pinterest.

Geo experiments skip user-level targeting. Comparable regions get split into test and control geographies, ad spend changes in the test markets, and the resulting difference in outcomes across the two sets of geos becomes the estimate of incremental lift. Because the whole thing is aggregated at the market level, geo tests are essentially immune to the signal loss caused by ATT and cookie deprecation, and they work across every channel at once, not just paid social. The cost isn't PSA spend, it's the opportunity cost of dark(ened) control markets, and the timeline tends to run 4 to 8 weeks or longer once a proper pre-period is included. Geo tests also need real scale: results get noisy below a certain spend threshold, and matching markets closely enough to trust the comparison isn't trivial. Meta has published its GeoLift framework and power-simulation tooling openly, which has made this design more accessible than it used to be. Availability elsewhere spans TikTok's Conversion Lift Study, Snapchat's managed lift programs, Pinterest's GeoX offering through partners alongside its Conversions API, and Google's geo and conversion lift tools documented in its Experiments Playbook.

For most DTC and ecommerce operators, geo-holdout testing is the workhorse: it's auditable, it doesn't depend on any single platform's internal logic, and it doesn't require a statistics background to interpret directionally. Ghost ads earn their keep in a narrower but important case, high-volume retargeting programs where CPM variance meaningfully swings cost per lead and speed of the read matters more than cross-channel visibility.

Choosing the right design for your budget, volume, and question

The question being asked should decide the design, not the other way around. Testing whether one channel is incremental is a different problem than testing whether the entire media mix is pulling its weight, which is different again from testing whether a specific campaign is bringing in genuinely new buyers.

User-level holdouts, whether ghost ads or a platform's native lift product, are where most DTC brands should start. They work at nearly any spend level for single-channel questions and retargeting validation, and the platform handles the suppression logic, which lowers the operational lift considerably. Geo holdouts are what brands graduate to once they're running meaningful spend across several channels simultaneously, because a user-level test inside one platform simply can't see cross-channel or offline effects; a geo test can, since it measures the market, not the individual user.

Time-based on/off testing is what gets run when a team isn't ready to commit to a real design, and it should be treated that way: a rough sanity check on a single channel, never the basis for a strategic budget call. Matched-market testing tied to one model is the other end of the spectrum, reserved for large brands running that kind of modeling on a quarterly cadence with enough historical data to make the modeling investment worthwhile. It's not a first test for anyone.

Between ghost ads and PSA holdouts specifically, the choice often comes down to infrastructure. Ghost ads make more sense when spend flows through a demand-side platform across exchanges, or when heavy mid-flight algorithmic optimization would otherwise contaminate a PSA control group, or simply when the budget cost of running PSAs is a real constraint. PSA holdouts become the fallback when ghost-bidding infrastructure isn't available on the platform in question, which is still common outside of Meta and a handful of DSPs.

Platform-native conversion lift products are fast to deploy, but they carry a structural bias: Meta has a financial incentive to report that Meta spend is working, and that ceiling should shape how much weight any single-platform result gets in a broader decision. This matters even more for omnichannel brands selling across Amazon, TikTok Shop, and retail simultaneously. A result from Meta's own Incremental Attribution tooling only measures what Meta can see. It isn't wrong, it's incomplete, and treating an incomplete number as the full picture is how budgets get misallocated across channels the platform has no visibility into.

Sizing the control group: the trade-off between statistical power and revenue left on the table

Every holdout test runs into the same tension. A bigger control group produces a more statistically confident result, but every user sitting in that control group is a user who didn't get the chance to convert off the ad, and that's real revenue withheld for the duration of the test. Many teams default to a 10 to 20% holdout because it sounds reasonable, but the right number depends entirely on how many conversions the brand generates in a given week, not on convention.

A practical sizing framework, drawn from adlibrary.com (May 2026), lays this out by conversion volume. Small brands generating fewer than 50 weekly conversions need a holdout of at least 50%, a minimum test duration of 8 to 12 weeks, and even then can only reliably detect a lift of 25% or more. Mid-tier brands at 50 to 500 weekly conversions can drop the holdout to 20 to 30%, run for 4 to 6 weeks, and detect lift down to around 15%. Large brands at 500 to 2,000 weekly conversions need only a 10 to 15% holdout over 2 to 4 weeks to detect an 8% lift. Enterprise-scale brands above 2,000 weekly conversions can run a 5 to 10% holdout for as little as 1 to 2 weeks and still detect lift as small as 5%.

Meta's own documentation, cited in that same adlibrary.com sourcing, specifies that Conversion Lift needs at least 100 incremental conversions to reach statistical significance on a conversion event, with practical guidance pointing toward a minimum run time of two weeks. Below that volume, the platform's own tool is telling advertisers the result won't be trustworthy.

The rule most teams get backwards is that minimum detectable effect should set the sample size, not the reverse. Deciding in advance how small a lift matters to the business, say 10% versus 30%, determines how many observations the test needs; running the test first and hoping the traffic happens to be enough is how noisy, unusable results get produced. If the available traffic volume can't support detecting the lift that actually matters to the business, the honest move is not to run the test at all, because it will generate noise dressed up as an answer.

A 10% holdout over a four-week test means withholding ads from one in ten prospective customers for a month, and whatever incremental lift rate the campaign would have produced against that slice of revenue is the opportunity cost. A 10% holdout over a four-week test means withholding ads from one in ten prospective customers for a month, and whatever incremental lift rate the campaign would have produced against that slice of revenue is money left on the table for that period. Most DTC brands running these tests conclude the data is worth that short-term cost, since a bad budget decision made on inflated attribution numbers costs far more over a full quarter. There's no universal "correct" holdout percentage to memorize here; it flows entirely from the conversion volume on hand and the minimum lift the business actually cares about detecting.

Setting test duration: why ending early is the most common and most costly mistake

Cutting a test short is, by a wide margin, the single most cited failure mode across sources on incrementality testing. It happens for mundane reasons: the early numbers look encouraging, budget owners get impatient, and nobody stops to ask whether a five-day read carries enough statistical weight to mean anything.

Minimum durations vary by design. Ghost ads and platform-native lift tools can sometimes deliver a valid read in 2 to 4 weeks if conversion volume is sufficient, but that should be confirmed with the platform's own power-calculation tools before launch, not assumed. PSA holdouts typically run a similar range, a few weeks, again scaled to conversion volume and however the control split was set. Geo experiments run longest, usually 4 to 8 weeks or more once the pre-period is factored in; GeoLift's own methodology stresses pre-period planning specifically because a rushed or thin pre-period produces an underpowered test regardless of how long the actual experiment runs.

Duration matters for reasons beyond raw statistical power, too. Day-of-week purchase patterns, weekly promotional cadences, and payday cycles all introduce noise that only smooths out over multiple cycles. A fourteen-day test that spans two full weekends captures a meaningfully more representative slice of behavior than a ten-day test that happens to miss one of them.

Short tests are also far more exposed to outside shocks: a competitor launching a surprise promotion, a platform algorithm update, a major news event, an outage. Any one of these can swing a short test's result dramatically, while a longer test simply absorbs the noise as one data point among many. If someone assigned to a control geo sees the ad anyway on a connected device, crosses into a treatment geo, or shares a household with someone in the treatment group, the control stops being clean. Longer tests and larger geo groupings reduce that risk, though they don't erase it.

The practical fix is straightforward. Pull the starting duration from the sizing table based on conversion volume, then check that it spans at least two full purchase cycles for the category in question. Whichever number is longer wins; shaving days off either one to hit a reporting deadline defeats the purpose of running the test.

Timing the test against seasonality and platform events

Even a well-sized, well-designed holdout test can produce a misleading result if it runs during the wrong window. The time-based on/off test fails partly for this exact reason, since organic demand doesn't pause for the experiment, but a properly randomized holdout isn't immune to timing distortion either; it's just distorted in a subtler way.

Certain windows should be avoided. Q4's peak stretch, from Black Friday/Cyber Monday through mid-December, pushes conversion rates well above baseline in ways that won't generalize to the rest of the year, so a lift measured there tells a brand about Q4, not about its media strategy broadly. Running a holdout during an active site-wide promotion creates the same problem in a different form: it becomes nearly impossible to separate the incremental effect of the sale from the incremental effect of the ads promoting it. The first week or two after a major platform algorithm or policy change is another bad window, since CPMs and delivery patterns are still settling into a new equilibrium. New campaign flights and brand launches deserve the same caution, since the algorithm needs time to stabilize delivery before a measurement is trustworthy.

Cleaner windows tend to be the steady-state stretches where organic conversion behavior is stable and boring in the best sense; mid-January through March and May through early August are commonly cited as reliable periods for most DTC categories. For geo experiments specifically, the pre-period isn't optional scaffolding, it's the baseline the whole synthetic control is built from, and GeoLift methodology treats a longer, cleaner pre-period as directly responsible for a more precise final read.

Timing discipline pays off concretely when it comes to branded search. A published dataset covering 225 DTC geo-based incrementality tests run between August 2024 and December 2025 (via stellaheystella.com) found a median iROAS on branded search of just 0.70x, well under the 1.0x breakeven point. That's a category where testing in a clean, uncontaminated window matters enormously, because a brand that draws budget conclusions about branded search from a noisy or seasonally skewed test risks either overfunding a line item that's actually losing money, or cutting one that looks worse than it is.

Reading iROAS and the incrementality gap honestly

The output that matters at the end of all this design work is incremental ROAS: incremental revenue divided by ad spend. It is not the same figure as the platform-reported ROAS in the dashboard, and it will almost always come in lower, sometimes dramatically so.

The branded search figure from the 225-test dataset (digitalapplied.com, covering August 2024 through December 2025) puts this in stark relief: a median iROAS of 0.70x means the median brand tested was, in incremental terms, losing money on branded search spend relative to what it would have captured organically anyway. That's not a fringe result buried in a footnote; it's the median across a couple hundred real DTC tests. A brand seeing an ugly number like that has two honest paths: cut the spend and redirect it toward channels with a proven incremental lift, or accept the spend as a defensive move (denying a competitor that search real estate) with eyes open about what it's actually buying. What isn't defensible is continuing to report the platform-attributed ROAS on branded search as if it reflects the same thing the holdout test just measured. Once the gap between attributed and incremental is visible, going back to the inflated number is a choice, not an oversight, and it's the kind of choice a holdout test exists specifically to make impossible to ignore.

Sources

  1. Incrementality Testing in 2026: Measure Causal Ad Lift
  2. Holdout Test in 2026: Incrementality Measurement That Actually Works
  3. Implementing holdout and ghost ads step by step - Customer Science
  4. Incrementality Testing: Proving Ads Actually Caused Sales

More in Incrementality Testing