Are Your B2B Ads Creating Demand or Taking Credit for It? An Incrementality Testing Guide

B2B marketing incrementality testing with account and geographic holdouts, The Geisheker Group, Inc.

Incrementality testing is a controlled experiment that withholds advertising from a randomly assigned group of accounts or geographies, then measures the difference in pipeline between the group that saw the advertising and the group that did not. Attribution answers which touchpoint deserves the credit for a deal that already happened; incrementality answers whether that deal would have happened anyway. Peter Geisheker, founder of The Geisheker Group, Inc., puts the problem plainly: “I have tested thousands of ads. I have never once correctly picked the winner. Not one time.”

Key Facts at a Glance

  • Attribution assigns credit. Incrementality tests causation. A channel can be credited with every deal it touched and still be creating none of them.
  • When eBay switched off brand-keyword paid search in 68 of 210 US media markets for 60 days, 99.5% of the forgone click traffic was immediately recaptured by natural search (Blake, Nosko and Tadelis, Econometrica, January 2015; experiment run on eBay’s own spend of roughly $51 million per year in US paid search).
  • On the same eBay data, the standard observational method estimated a 4,100% return on non-brand search; the controlled experiment estimated minus 63% (Blake, Nosko and Tadelis, Econometrica, January 2015).
  • Across 15 large-scale advertising randomized controlled trials, observational methods were off by a factor of three in half the studies, and usually in the direction of overstating the advertising’s effect (Gordon, Zettelmeyer, Bhargava and Chapsky, Marketing Science, March 2019; 500 million user-experiment observations and 1.6 billion impressions, conducted on Facebook data).
  • In 25 large retail and brokerage field experiments, the median standard error on measured ROI was 26.1%, producing a confidence interval more than 100 percentage points wide (Lewis and Rao, Quarterly Journal of Economics, November 2015).
  • Peter Geisheker’s operating position on why most tests answer the wrong question: “The ad buys the click. The landing page buys the lead. Most companies test one and pray about the other.”
  • A B2B test window shorter than the sales cycle cannot see the result it was built to measure. Peter Geisheker: “If the sales cycle is six months, the fractional CMO hired today has no sales to show for six-plus months.”

Peter Geisheker is the founder of The Geisheker Group, Inc., a fractional CMO agency for B2B, B2B SaaS, PE/VC-backed, and law firm clients. He has managed more than $50 million in advertising spend across a 20-year career built on direct response, where the entire discipline rests on the question this article is about: did the advertising cause the sale, or did it merely stand next to it.

Table of Contents

Need marketing leadership and expert strategy to grow your company?

The Geisheker Group is a B2B fractional CMO agency that installs the measurement systems most companies never get around to building.

Explore Fractional CMO Services

What is incrementality testing, and how is it different from attribution?

Incrementality testing is a controlled experiment that withholds advertising from a randomly assigned group of accounts or geographies, then compares pipeline between the group that saw the advertising and the group that did not. Attribution and incrementality answer different questions. Attribution asks which touchpoints deserve credit for deals that already closed. Incrementality asks whether those deals would have closed without the advertising at all.

The distinction matters because attribution is structurally incapable of finding a zero. Every attribution model, first touch, last touch, linear, time decay, or a custom B2B lead attribution model, starts from a population of people who converted and then distributes credit among the touchpoints those people encountered. The math has no mechanism for concluding that a touchpoint created nothing. If a buyer who was going to sign anyway happened to click a retargeting ad on the way, that ad receives credit, and the credit looks identical to credit earned by an ad that changed someone’s mind.

Incrementality inverts the question. Instead of asking who was present at the conversion, it asks what happens to conversions when you remove the advertising. The only way to answer that is to remove it from some people and not others, then compare.

Attribution Incrementality
Question answered Which touchpoint gets credit for this deal? Would this deal have happened without the advertising?
Method Observational; models credit across the touchpoints converters encountered Experimental; withholds advertising from a randomized control group
Can it return zero? No. Every model distributes 100% of credit among observed touchpoints Yes. A result of no measurable lift is a valid and common outcome
Main failure mode Credits advertising that reached buyers who were already converting Requires enough units and enough time to detect an effect, which many B2B programs do not have
What it is good for Budget allocation between channels that are all known to work Deciding whether a channel works at all
Cost of being wrong You shift budget toward the channel best at being present You shut off a channel that was working, or keep one that was not

Both have a place. Attribution is a reasonable way to divide budget among channels you have already established are productive. It is a catastrophic way to decide whether a channel is productive in the first place.

Why do attribution reports overstate what your advertising is doing?

Attribution reports overstate advertising because advertising is deliberately aimed at the people most likely to buy, and those people were the most likely to buy before any advertising reached them. Every targeting system, every retargeting pool, and every algorithmic bidding engine is built to find high-intent buyers. The better the targeting, the more the advertising overlaps with demand that already existed, and the larger the overstatement becomes.

The strongest public demonstration of this comes from eBay. Researchers Thomas Blake, Chris Nosko and Steven Tadelis ran a controlled experiment on eBay’s own paid search, published in Econometrica in January 2015. When the company stopped bidding on its own brand keywords across 68 of the 210 US designated market areas while 142 served as controls, 99.5% of the forgone click traffic was immediately recaptured by natural search results. The clicks that paid search had been billing for were clicks eBay was going to receive regardless.

The non-brand result is the one that should worry any marketer reading a dashboard. Over a 60-day test, the experiment estimated a return on investment of minus 63%, with a 95% confidence interval running from minus 124% to minus 3%. The same data analyzed the way a normal attribution report analyzes it, a regression without experimental controls, produced an estimated ROI of 4,100%. Adding geographic and time controls, which is more rigor than most marketing teams apply to anything, still produced 1,400%.

This is not an eBay quirk. Brett Gordon, Florian Zettelmeyer, Neha Bhargava and Dan Chapsky compared observational measurement against randomized controlled trials across 15 large-scale advertising experiments covering 500 million user-experiment observations and 1.6 billion impressions, published in Marketing Science in March 2019. In half the studies, the estimated percentage increase in purchase outcomes was off by a factor of three across all the observational methods they tried. In one study the true experimental lift on checkouts was 73%, while the simple exposed-versus-unexposed comparison reported 316%. The bias generally ran toward overstating advertising’s effect, though not always, and no single observational method was reliably better than the others.

The practical conclusion is uncomfortable and worth stating flatly. If your only evidence that a channel works is that the channel’s dashboard says it works, you do not have evidence. You have a measurement built on the same selection that the advertising itself performed.

How does an account holdout work in B2B?

An account holdout randomly assigns target accounts to a treatment group that receives the advertising and a control group that is suppressed from it, then compares pipeline creation between the two groups over a window at least as long as the sales cycle. It is the natural design for B2B because B2B spend is usually aimed at a defined list of accounts rather than at an open market, and because the account, not the individual, is the unit that generates revenue.

The mechanics, in order:

  1. Start from a fixed target list. Everything in the test must come from one defined universe, typically your ICP list or an ABM tier. Accounts that enter mid-test break the randomization.
  2. Randomize at the account level, not the contact level. Multiple people from one buying committee must land in the same arm. If three contacts at the same company split across treatment and control, the control contact hears about you from a colleague and your control is no longer a control.
  3. Stratify before you randomize. Rank accounts by the thing that predicts pipeline, usually revenue band, employee count, or prior engagement, then randomize within each band. Simple randomization on a list of 200 accounts will produce lopsided groups often enough to ruin the test.
  4. Suppress, do not merely stop targeting. Control accounts go into an exclusion list on every platform in the test. On LinkedIn this is a company-list exclusion; on Meta and Google it means excluding the matched customer list and any lookalike built from it.
  5. Measure the outcome that money cares about. Opportunities created and pipeline dollars, not clicks and not MQLs. An incrementality test measured on form fills tells you whether the ads produced form fills, which nobody was in doubt about.
  6. Read the result at the account level. Percentage of accounts creating an opportunity, treatment versus control, plus the confidence interval around that difference.

Account holdouts are the right design when you have a list you control, a platform that supports list exclusion, and enough accounts to detect a difference. That last condition is where most B2B tests quietly fail, and it is covered below.

When should you use a geographic holdout instead of an account holdout?

Use a geographic holdout when you cannot suppress advertising at the account level, when the channel does not accept exclusion lists, or when your demand is broad rather than list-based. A geo holdout partitions a market into geographic units, turns advertising off in a randomly selected subset, and compares outcomes between the on and off regions. Google researchers Jon Vaver and Jim Koehler formalized the method in 2011, using the 210 US designated market areas as the standard partition.

The design has two hard requirements, stated in the original Google paper: you must be able to serve advertising according to a geographic prescription with reasonable accuracy, and you must be able to track both spend and the response metric at the geographic level. If your CRM does not reliably capture account location, a geo test cannot be read no matter how well it is run.

One design detail from that paper is worth applying directly, because it is free. Grouping geographies by size before randomly assigning them to test and control, rather than assigning at random across the whole set, can reduce the width of the resulting confidence interval by 10% or more. For a B2B advertiser with limited volume, a 10% tighter interval can be the difference between a readable result and a shrug.

Account holdout Geographic holdout
Best for ABM, list-based paid social, targeted display Broad-market search, brand campaigns, offline media
Unit of assignment Company DMA, state, metro, or sales territory
Requires Platform exclusion lists and a stable target list Geo-level ad delivery and geo-level outcome tracking
Main contamination risk Contacts at one account split across arms; organic and outbound reaching control accounts Buyers crossing geographic boundaries; imprecise geo-targeting; national accounts with multi-site buying
Typical B2B weakness Too few accounts to reach statistical power Too few geos with meaningful B2B volume; enterprise deals concentrated in a handful of metros
Reads best on Opportunity creation rate per account Pipeline dollars per geo, indexed to a pre-test baseline

For most mid-market B2B companies the account holdout is the better instrument, because the target universe is genuinely a list. Geographic holdouts come into their own for private equity portfolio companies with regional sales territories, where a territory is already the natural unit of both spend and revenue reporting.

Find out which half of your spend is creating business, and which half is documenting it.

Peter Geisheker has managed more than $50 million in advertising spend and delivered 6X inbound lead growth, 100% year-over-year SaaS revenue growth for three consecutive years, and a 77% reduction in paid acquisition costs. The Geisheker Group builds revenue systems that report honestly, including when the honest report is that a channel should be switched off.

See Fractional CMO Services

What contaminates a B2B incrementality test?

Contamination is anything that lets the control group receive the treatment, or that makes the two groups differ for a reason other than the advertising. B2B is unusually vulnerable to it, because B2B demand generation is deliberately multi-threaded: the same account is being reached by paid media, outbound sequences, events, organic search, and a salesperson, often in the same week. Every one of those threads is a path by which advertising you thought you withheld arrives anyway.

The most important contamination source is the one most teams never consider, because it was invented by the platforms themselves. Garrett Johnson, Randall Lewis and Elmar Nubbemeyer documented it in Journal of Marketing Research in 2017: on an algorithmically optimized ad platform, a public service announcement control group is no longer a valid control, because the platform’s delivery system chooses different users for the PSA than it would have chosen for your real ad. The control group and the treatment group stop being comparable populations. Their proposed fix, ghost ads, works by tagging the impressions where your ad would have won the auction for control users, so the comparison is between people the system actually selected for your campaign in both arms. The same paper found that designs relying only on intent-to-treat estimates carried variance 5.9 to 16.4 times higher, meaning such experiments need to be roughly an order of magnitude larger to reach the same confidence.

The practical register for a B2B test:

Contamination path What it does to the test Control
Contacts from one account split across arms Control accounts receive the advertising through a colleague Randomize at the account level and enforce it in every platform’s list
Lookalike and similar-audience expansion The platform serves your control accounts anyway, because they resemble converters Exclude the customer list and every audience derived from it, then verify delivery by account
Outbound and SDR sequences running in parallel Control accounts get touched by a different part of your own company Freeze outbound to both arms for the test window, or randomize outbound with the same assignment
Organic search and AI-generated answers Control accounts find you without advertising Cannot be removed; must be acknowledged as a floor under the control group, which makes measured lift conservative
Sales team working the target list Reps unknowingly prioritize accounts they saw engagement from Blind the sales team to arm assignment, and measure whether rep activity differs between arms
Mid-test list changes New accounts join, breaking the randomization Lock the universe before the test starts; log any change as a protocol deviation
Multi-site enterprise accounts in a geo test One buying unit spans test and control regions Assign at the parent-account level, or exclude multi-site accounts from a geo test
PSA-style control creative on an optimized platform The platform delivers control creative to a different population Use platform-native lift tooling with a true holdout, not a substituted creative

Note the organic search row carefully, because it cuts in a useful direction. Contamination that leaks demand generation into the control group biases the measured lift downward. If you measure real lift in spite of it, the finding is stronger than it looks. It is contamination running the other way, where the treatment group gets extra advantages the control does not, that manufactures fake wins.

How do you measure incrementality when the sales cycle is six months?

You extend the measurement window to cover the full sales cycle plus reporting lag, and you read interim results only on leading indicators that you have separately validated as predictive. A B2B incrementality test read at four weeks on a six-month sales cycle is not an early result. It is a measurement of a period during which almost none of the effect could have appeared yet, and treating it as a preview is how good programs get killed.

This is the structural problem underneath most B2B marketing measurement, and Peter Geisheker states it directly:

Peter Geisheker, founder of The Geisheker Group, Inc., frames it as a measurement-window failure rather than a marketing failure: “In a business with a long sales cycle, the marketing work you did six months ago is what shows up as won deals today, and that contradicts the quarterly number. If the sales cycle is six months, the fractional CMO hired today has no sales to show for six-plus months. Business is run on numbers, and marketing has to show how its numbers, short term AND long term, are moving the business, or finance measures you on a window shorter than your own sales cycle and punishes the work that’s actually working.”

The same mismatch that punishes a new marketing leader also invalidates a short incrementality test. Finance wants the read this quarter; the causal effect does not finish arriving until next quarter or the one after. There are three honest ways to handle it, and one dishonest one.

Run the test for the full cycle and say so up front. The cleanest option. Define the window as median sales cycle plus the lag between opportunity creation and CRM entry, get agreement on that number before launch, and commit to not reading the result early. A six-month cycle means a test that reports in month seven or eight.

Read interim results on a validated leading indicator, not on revenue. If you can demonstrate from historical data that opportunity creation at stage two predicts closed-won at a stable rate, then stage-two opportunity creation is a legitimate interim metric. The word doing the work there is demonstrate. Most teams assume the relationship rather than checking it, and an unvalidated leading indicator is just a shorter test with the same problem.

Test the front of the funnel instead, and be explicit about what you have proven. Measuring incremental lift on qualified opportunity creation is faster and far more achievable than measuring incremental lift on closed revenue. It is also a smaller claim. Proving a channel creates incremental opportunities does not prove it creates incremental revenue, because the incremental opportunities may convert worse than the organic ones. State that limit rather than letting the reader assume it away.

The dishonest option, which is common, is to run a four-week test, find no significant effect, and conclude the channel does not work. On a long-cycle B2B business that conclusion is unsupported by the design that produced it. A null result from a test that ran shorter than the sales cycle is not evidence of no effect. It is an absence of evidence, and those are different things.

When is your sample too small to conclude anything?

Your sample is too small when the confidence interval around your measured lift is wide enough to contain both the outcome that would justify the spend and the outcome that would justify cancelling it. That is the operational test, and it is more useful than any rule of thumb about minimum counts, because it is specific to your own economics.

The underlying economics of ad measurement are brutal, and Randall Lewis and Justin Rao quantified them in the Quarterly Journal of Economics in November 2015 using 25 large field experiments with US retailers and brokerages. Their core finding: individual-level sales are so variable relative to advertising’s effect that detecting a real return requires enormous samples. In their data, mean sales per person over the campaign window was $7 with a standard deviation of $75. The median standard error on measured ROI across the retail experiments was 26.1%, which produces a confidence interval more than 100 percentage points wide. For the brokerage experiments the median standard error was 115%. To reliably distinguish a 50% ROI from a 0% ROI, they calculated the median campaign would have to be nine times larger; to resolve a 10-percentage-point difference in ROI, 62 times larger. Their summary line is that informative advertising experiments can easily require more than 10 million person-weeks.

Those experiments were consumer campaigns reaching millions of people. B2B does not have millions of accounts. This is the part of the incrementality conversation the vendor blogs skip, so state it plainly: for most B2B companies, a statistically conclusive incrementality test on closed revenue is not achievable at the account level, and no amount of methodology fixes that. The units do not exist.

Run the arithmetic on a realistic B2B design and the problem becomes concrete. Take 800 target accounts split evenly into treatment and control, with 8% of control accounts expected to create an opportunity during the test window. At 95% confidence and 80% power, the smallest lift that design can detect is a 67% relative increase in opportunity creation. Not 5%, not 20%. Sixty-seven percent. To detect a 20% lift with the same baseline rate you would need roughly 4,500 accounts in each arm, which is 9,000 accounts in a target universe most mid-market B2B companies do not have. The companion worksheet computes this for your own numbers.

What that implies is not that you should give up. It is that you should choose your test so the arithmetic can work:

  • Move the outcome metric earlier in the funnel. Opportunity creation happens at a far higher rate than closed-won, so the same account count produces many more events and a tighter interval. You learn something smaller, but you actually learn it.
  • Test bigger differences. A test designed to detect a 5% lift needs vastly more units than one designed to detect a 40% lift. If the channel is 30% of your paid budget, the decision-relevant question is usually “is it doing roughly what we think, or roughly nothing,” and that is a large-effect question.
  • Use geographies or territories as units when you have more of them than you have meaningful accounts. A company with 40 sales territories and 900 target accounts may get a cleaner read from 40 units of aggregated pipeline than from 900 units of mostly-zero account outcomes.
  • Extend duration rather than shrinking the effect you are hunting. More weeks of the same units accumulates events, though it also accumulates contamination and seasonality risk.
  • Run a directional pilot and label it as one. A test that cannot reach significance can still be worth running if the result is treated as a prior rather than a verdict, and if nobody writes “proven” in the deck.

Before launching, compute one number: given your baseline conversion rate, your unit count, and your test duration, what is the smallest lift you could detect at 80% power? If that minimum detectable effect is larger than any lift you would plausibly see, the test is already finished and the answer is that you cannot run it. Finding that out in the planning meeting costs an hour. Finding it out after a six-month test costs a budget cycle.

What goes in a test-planning worksheet?

A test-planning worksheet forces every decision that determines whether a result is readable to be made before the test starts, rather than negotiated afterward when the numbers are already in view. The most important rows are the ones teams skip: the minimum detectable effect, the contamination register, the decision rules written in advance, and an explicit statement of what the test will not be able to conclude.

Fill this in completely before a single campaign is paused. Any row you cannot answer is a reason to delay the test, not a reason to start it and hope.

# Field What goes here Why it matters
1 Hypothesis One falsifiable sentence: “Paid LinkedIn to tier-1 accounts creates incremental qualified opportunities.” A hypothesis you cannot disprove is not a test
2 Channel or campaign under test The exact campaigns, ad accounts, and audiences being withheld Ambiguity here means you cannot say what was actually tested
3 Unit of assignment Account, parent account, geography, or territory The unit must match how buying actually happens
4 Universe and lock date The defined list, its size, and the date it was frozen Accounts entering mid-test break randomization
5 Stratification variables Revenue band, employee count, prior engagement, industry Prevents lopsided groups on small samples
6 Holdout size and split Number and percentage of units in control Drives statistical power directly
7 Primary outcome metric One metric, named exactly as it appears in the CRM Multiple primary metrics invite cherry-picking
8 Secondary and guardrail metrics Cost per opportunity, sales-accepted rate, pipeline quality Catches a lift that is real but worthless
9 Pre-test baseline Outcome rate per unit per period over the prior 2 to 4 cycles Without it you cannot compute power or detect drift
10 Baseline variability Standard deviation of that outcome across units The denominator of every power calculation
11 Minimum detectable effect Smallest lift detectable at 80% power with this design If this exceeds any plausible lift, stop here
12 Sales cycle and reporting lag Median days to close plus days to CRM entry Sets the floor on test duration
13 Test window Start date, end date, and read date, with the lag built in Prevents reading the test before the effect can appear
14 Suppression checklist Every platform where control units must be excluded, verified after launch Suppression that was configured but not verified is the most common silent failure
15 Contamination register Each leak path, its control, and whether the control is in place Forces honesty about what cannot be controlled
16 Parallel activity freeze What outbound, events, and email are paused or held constant Your own company is the biggest contamination source
17 Interim read policy What may be looked at before the read date, and by whom Prevents a null four-week result from killing the test
18 Decision rules Written before launch: if lift exceeds X do A; if the interval crosses zero do B Decision rules written after the result are rationalizations
19 Stop conditions What would end the test early, such as a pipeline collapse in control Protects the business from the experiment
20 Audience and read date Who receives the result, in what format, on what date An unread test is a cancelled test with extra steps
21 What this test cannot conclude Written out explicitly, in plain language The single most valuable row on the sheet

Row 21 deserves the emphasis. Every test has a boundary, and writing it down before launch is what stops a narrow finding from being reported as a broad one. “This test can tell us whether paid LinkedIn creates incremental opportunities among tier-1 accounts over 90 days. It cannot tell us whether those opportunities close at the same rate as organic ones, and it cannot tell us anything about tier-2 and tier-3 accounts.”

A companion spreadsheet version of this worksheet, with the power-calculation inputs and the contamination register as working tabs, accompanies this article.

What do you do when the test says a channel is not incremental?

You cut the spend, redeploy it to something you have not yet disproven, and then test that. A non-incremental channel is not a failure of the marketing team; it is money that was buying demand that already existed, and finding it is the entire point of running the test. The uncomfortable part is that the channel’s dashboard will keep reporting conversions after you cut it, because those conversions were always going to happen.

Peter Geisheker has watched the shape of this in a live account, though it is worth being precise about what the episode does and does not prove:

Peter Geisheker, founder of The Geisheker Group, Inc., on rebuilding a B2B SaaS client’s paid search: “A SaaS client was bidding on broad, general search terms on Google. It was costing them a fortune and producing almost nothing. I rebuilt the campaigns around long-tail searches an actual buyer would type. More leads immediately, and cost per lead went from about six hundred dollars to around one hundred.”

That result was not produced by a controlled experiment, and it is worth being exact about what it can and cannot support. The generic campaigns were paused and the focused campaigns were built at the same time, with no concurrent control group, so the episode cannot isolate which change caused what. A large block of spend came out and lead volume did not collapse behind it, but “the generic spend was creating nothing” and “the focused campaigns replaced what the generic ones were creating” both fit those facts, and nothing in the account distinguishes them.

That is not a weakness in the example. It is the example. Most of what marketing teams call evidence has exactly this structure: a change was made, the numbers moved in a welcome direction, and the causal story gets written afterward by whoever is most confident. The reason to run a real holdout is that it is the only way to tell those two readings apart, and the reason to write down what your test cannot conclude is that without it you will reach for the flattering reading every time.

What the episode does give you is the shape worth watching for, and that generalizes:

  1. Cut, then watch the total, not the channel. The question is never whether the channel’s reported conversions fell, because they will. It is whether total qualified pipeline fell.
  2. Wait one full sales cycle before concluding the cut was safe. A cut looks free for exactly as long as the pipeline built before the cut keeps converting.
  3. Redeploy rather than bank. Moving the spend to an untested channel and testing that one converts a dead channel into information. Returning the money to finance converts it into a smaller budget next year.
  4. Re-test on a schedule. A channel that is non-incremental at your current brand awareness and offer may become incremental after either changes. Non-incremental is a finding about now, not forever.
  5. Separate the channel from the execution. A channel can fail to produce lift because the channel cannot work for you, or because what you ran in it was bad. As Peter Geisheker puts it: “The ad buys the click. The landing page buys the lead. Most companies test one and pray about the other.” Before concluding a channel is dead, confirm that the destination was not the thing that killed it.

The broader discipline here is the same one that separates measurable advertising from the unmeasurable kind, and it is the reason attribution alone cannot settle these questions. Credit assignment describes the past. Incrementality is the only tool that tells you what to do next.

Frequently asked questions

What is incrementality testing in B2B marketing?

Incrementality testing in B2B marketing is a controlled experiment that withholds advertising from a randomly assigned group of accounts or geographies and compares pipeline creation against the group that received it. Unlike attribution, which distributes credit among the touchpoints a converting buyer encountered, incrementality measures whether the advertising caused business that would not otherwise have occurred.

What is the difference between incrementality and attribution?

Attribution assigns credit for deals that already closed; incrementality tests whether those deals were caused by the advertising. Attribution is observational and cannot produce a result of zero, because every model distributes 100% of credit among observed touchpoints. Incrementality is experimental and returns zero regularly, which is precisely what makes it useful for deciding whether to keep funding a channel.

How long should a B2B incrementality test run?

A B2B incrementality test should run for at least the median sales cycle plus the lag between opportunity creation and CRM entry. For a six-month sales cycle, that means reading results in month seven or eight. Reading a long-cycle test at four weeks measures a period during which almost none of the causal effect could have appeared, and a null result from such a test is not evidence that the channel does not work.

How many accounts do you need for an account holdout test?

There is no universal minimum; the requirement depends on your baseline conversion rate, the size of the lift you need to detect, and your test duration. The correct approach is to compute the minimum detectable effect at 80% power before launching. Research by Randall Lewis and Justin Rao in the Quarterly Journal of Economics found that detecting economically meaningful returns often requires samples far larger than a single campaign provides, and B2B account lists are orders of magnitude smaller than the consumer campaigns they studied.

What is a geo holdout test?

A geo holdout test partitions a market into geographic units such as designated market areas, states, or sales territories, randomly switches advertising off in a subset, and compares outcomes between the on and off regions. Google researchers Jon Vaver and Jim Koehler formalized the method in 2011. It requires the ability to serve advertising by geography and to track both spend and the outcome metric at the geographic level.

What contaminates an incrementality test?

Contamination is anything that lets the control group receive the treatment or makes the groups differ for reasons other than the advertising. In B2B the most common sources are contacts from one account landing in both arms, lookalike audiences serving control accounts anyway, parallel outbound sequences, sales reps unknowingly prioritizing engaged accounts, and organic search reaching control accounts. Research published in Journal of Marketing Research in 2017 also established that public service announcement control groups are invalid on algorithmically optimized platforms, because the delivery system selects different users for the control creative.

Can you test incrementality on brand advertising?

Yes, using geographic holdouts, though the sample size and duration requirements are substantially harder than for direct response. The more practical alternative for most B2B companies is to add a response mechanism to brand advertising so it produces measurable signal continuously rather than requiring a dedicated experiment. Peter Geisheker’s standing position is that every placement can carry at least a lead magnet, which makes the advertising testable without waiting for a formal lift study.

What should you do if a test shows no measurable lift?

First check whether the test could have detected a lift at all, by comparing your minimum detectable effect against the lift you were looking for. If the design was underpowered, the result is inconclusive rather than negative. If the design was adequate, cut the spend, move it to an untested channel, watch total qualified pipeline rather than the channel’s own reporting for one full sales cycle, and re-test the cut channel later, since non-incremental is a finding about current conditions rather than a permanent verdict.

Installing this in your company

Understanding incrementality is straightforward. Installing it is not, and the reason has almost nothing to do with statistics. A real holdout means withholding marketing from accounts your sales team wants covered, holding the line when a four-week interim number looks bad, and committing in advance to a decision rule that may require cutting a channel someone built their year around. Those are organizational problems wearing a measurement costume. The math is the easy part.

That installation is fractional CMO work with a specific scope: defining the test universe, negotiating the outbound freeze with sales leadership, computing the minimum detectable effect before anyone commits budget, writing the decision rules while the outcome is still unknown, and reporting the result honestly to a CEO or board hoping for a different number. It takes someone with the standing to tell an executive that a channel they like is buying demand that already existed.

This is not the right project for every company. If your account list is too small for any design to reach statistical power, tighten conversion tracking and lead qualification first instead. If your marketing spend is under roughly $20,000 a month, the test can cost more than the waste it finds. And if the organization will not act on a null result, the test is theater.

If none of those describe you, schedule a 30-minute call. We will look at your account volume, sales cycle, and current measurement, and tell you honestly whether a valid test is achievable at your scale.

About Peter Geisheker

Peter Geisheker is the founder and CEO of The Geisheker Group, Inc., a fractional CMO agency serving B2B, B2B SaaS, private-equity-backed, and law firm clients. Over more than 20 years in direct-response marketing he has managed over $50 million in advertising spend, delivered 6X inbound lead growth, driven 100% year-over-year SaaS revenue growth for three consecutive years, reduced paid acquisition costs by 77%, and run programs deploying up to $1 million per week. His work centers on building capital-efficient revenue systems that report honestly, including when the honest report is that spend should stop.

More about Peter Geisheker and on LinkedIn.

References and sources

  1. Thomas Blake, Chris Nosko and Steven Tadelis, “Consumer Heterogeneity and Paid Search Effectiveness: A Large-Scale Field Experiment,” Econometrica, Vol. 83, No. 1, January 2015. Author copy, full text.
  2. Brett R. Gordon, Florian Zettelmeyer, Neha Bhargava and Dan Chapsky, “A Comparison of Approaches to Advertising Measurement: Evidence from Big Field Experiments at Facebook,” Marketing Science, Vol. 38, No. 2, March 2019. Full text.
  3. Randall A. Lewis and Justin M. Rao, “The Unfavorable Economics of Measuring the Returns to Advertising,” The Quarterly Journal of Economics, Vol. 130, No. 4, November 2015. Full text.
  4. Garrett A. Johnson, Randall A. Lewis and Elmar I. Nubbemeyer, “Ghost Ads: Improving the Economics of Measuring Online Ad Effectiveness,” Journal of Marketing Research, Vol. 54, No. 6, 2017. Full text.
  5. Jon Vaver and Jim Koehler, “Measuring Ad Effectiveness Using Geo Experiments,” Google Inc., 2011. Full text.

Similar Posts