Incrementality testing: is your dental marketing causing the bookings?
On this page
Incrementality testing for dental marketing asks a narrower question than any report does: how many bookings would you have lost if the advertising had not run? You answer it by withholding the ads from a comparable group of people, areas or weeks, then comparing their bookings with the group that saw the ads. The difference is what the marketing caused.
It's the most direct answer to "does my dental advertising actually work?", and most single practices don't have the volume to get it cleanly.
Attribution is not causation
Attribution assigns credit for a booking to the marketing that came before it. How the models split that credit is covered in how to measure dental marketing. None of them can see what a patient would have done without the ad, because they only look at patients who did book.
Take a patient who searches your practice name, clicks your brand ad and books. Attribution credits the ad. But your practice was probably the first organic result for that search too, and without the ad the patient may well have clicked it and booked anyway. Whether they would have is the incremental question, and no report built from clicks can answer it.
The gap matters most for spend that sits close to the booking, such as brand search and remarketing. It looks strong in attribution reports because it's present when people who had already decided book. Even with dental marketing attribution joined from click to completed treatment, this question stays open.
Holdout and geo tests: how incrementality testing works
Every incrementality test has the same shape: split the audience into two comparable groups, show the marketing to one, withhold it from the other, and compare. A holdout is the group that doesn't see the ads. Random assignment makes the groups comparable; splitting by area, a geo experiment, is the practical version when random assignment isn't available.
| Design | How it works | Who can run it | Main weakness |
|---|---|---|---|
| Platform user holdout | Platform randomly withholds ads from a control group | Google and Meta Conversion Lift, if you qualify | Counts platform conversions, which may not be bookings |
| Platform geo study | Platform withholds ads in some regions | Google Conversion Lift based on geography, via a Google representative | Not self-serve |
| Your own geo holdout | Exclude some areas in location settings; count bookings by patient postcode | Anyone with clean postcode data | Few areas in a catchment, and people travel between them |
| On and off in time | Pause a campaign for some weeks and compare | Anyone | Seasonality and everything else that changes |
| Google Ads custom experiment | Splits traffic between a campaign and a changed copy | Self-serve | Tests a change, not whether the campaign adds bookings |
What the platforms offer, checked 1 October 2026
Google Ads. Google describes Conversion Lift as an incrementality tool comparing people who could see your ads with a control group who could not, in user-based and geography-based versions. It isn't available for all accounts; you contact a Google account representative to use it.1 The user-based version requires a minimum campaign budget of $5,000 USD and at least 1,000 observed conversions.2 The geography-based version needs campaigns that target a single country and can use offline data; Google's page states no budget minimum, so I treat its threshold as unverified.3
Custom experiments are self-serve on Search, Display and Video campaigns, among others.4 They split traffic between your campaign and a modified copy, so there's never a group without ads.5 For Performance Max, Google runs uplift experiments that measure what adding Performance Max does alongside comparable campaigns, available for some goals including lead generation.6
A do-it-yourself geo holdout needs no special access: you exclude the holdout areas in the campaign's location settings. Google notes its location targeting doesn't follow exact borders, so some people in holdout areas will still see ads.7
Meta. Meta's Conversion Lift also compares a test group with a control group that can't see your ads. As a guide, Meta says the ad account needs a campaign that started in the past year with $5,000 USD or more of spend and at least 500 conversions, plus a signal-quality requirement such as Conversions API with an event match quality score above 5.8 It recommends at least $5,000 USD of budget and 28 days.9 Meta says one signal route is unavailable to advertisers mainly targeting countries affected by the EU ePrivacy Directive.8 I have not confirmed how Meta treats UK-targeted accounts, so that point is unverified. For a practice, the signal requirement also runs into the limits on sending health-related events to Meta, set out on the Meta Conversions API page.
How much volume a test needs
Bookings arrive unevenly even when nothing changes. The standard deviation of a count of independent events is roughly its square root.10 Illustrative figures: a diary that averages 30 consultation bookings a month will typically move by about 5 either way from month to month with no change at all, and a real diary can be noisier than that because of holidays, clinician leave and full weeks.
A test has to find the ads' effect inside that noise. A workable rule of thumb: the gap between test group and holdout needs to be about twice its own noise, and the noise in a gap between two counts is roughly the square root of their sum. The table applies that rule to hypothetical effects.
| Hypothetical effect (illustrative only, not a benchmark) | Holdout bookings needed, bare minimum | Holdout bookings needed, comfortable |
|---|---|---|
| Test group books 2 times the holdout | 12 | 24 |
| Test group books 1.5 times the holdout | 40 | 78 |
| Test group books 1.25 times the holdout | 144 | 282 |
| Test group books 1.1 times the holdout | 840 | 1,646 |
At the bare minimum, a real effect of that size is detected only about half the time; the comfortable column detects it most of the time. This is the sample size question, and modest effects, the ones budget decisions turn on, need hundreds of bookings per group. Count bookings from your practice management system (PMS) by patient postcode; completed treatment usually arrives too late to read inside a test.
Reading a result
A worked example, with illustrative figures, not benchmarks. A practice splits its catchment into two matched sets of areas. Both run ads for eight weeks, then ads run only in set A for the next eight.
| Illustrative figures, not benchmarks | Set A (ads on) | Set B (holdout) |
|---|---|---|
| Consultation bookings, 8 weeks before | 24 | 23 |
| Consultation bookings, 8 weeks during | 30 | 22 |
| Change | +6 | −1 |
Comparing each set with its own baseline, then comparing the changes, gives an estimate of about 7 extra bookings from the ads in set A. Now the noise: the four counts add up to 99, the square root is about 10, and twice that is about 20. The plausible effect runs from roughly 13 fewer bookings to 27 more. The test can't show the ads did nothing, and it can't show they did something. The honest report is "inconclusive", not "7 extra bookings".
Google reports the certainty of lift from 50% to 95% and treats 90% or more as most reliable, with lower levels directional only.2 Meta describes a 90% or 95% threshold as the conventional test for whether an effect exists at all.11
Habits that keep the reading straight:
- Report the range, not the single number.
- Inconclusive isn't zero. It means the test was too small to tell.
- List what else changed: a new clinician, prices, reviews, a competitor, school holidays.
- Check capacity. A full diary can't show extra bookings, whatever demand the ads created.
- Repeat or extend a test before cutting spend on the strength of it.
A clear result belongs at the top of any calculation of whether the marketing paid for itself: incremental bookings, not attributed ones, are the right starting point.
When a practice should not test
Do not test when:
- Measurement isn't joined yet. Without each new patient's source and postcode recorded, you can't count bookings by group.
- The volume isn't there. Check your monthly bookings for the treatment against the table above.
- The holdout costs too much. Withholding emergency ads from half your catchment for eight weeks sends patients in pain elsewhere.
- The diary is already full. A full diary hides any effect.
- Something else is changing. A refit, a new associate or a new website will swamp the result.
When a test is out of reach, change one thing at a time, give it long enough to cover the time from enquiry to booking, and watch the whole appointment book rather than the channel's own report. Most of the levers in running Google Ads for a practice, such as brand bidding, location settings and budget, can be changed and read that way.
When the numbers don't support a test, I say so. If you want your own volumes checked against the table above, email me at [email protected]. The first look is free, done from the outside, and I reply with what I find. You can also read how I work first. Terms used here are defined in the dental marketing glossary.
Sources
-
Google Ads Help: About Conversion Lift, accessed 1 October 2026. ↩
-
Google Ads Help: Set up Conversion Lift based on users, accessed 1 October 2026. ↩ ↩2
-
Google Ads Help: Set up Conversion Lift based on geography, accessed 1 October 2026. ↩
-
Google Ads Help: About custom experiments, accessed 1 October 2026. ↩
-
Google Ads Help: Set up a custom experiment, accessed 1 October 2026. ↩
-
Google Ads Help: Experiments FAQs, accessed 1 October 2026. ↩
-
Google Ads Help: Exclude ads from geographic locations, accessed 1 October 2026. ↩
-
Meta Business Help Centre: About conversion lift, accessed 1 October 2026. ↩ ↩2
-
Meta Business Help Centre: How to set up a conversion lift test, accessed 1 October 2026. ↩
-
NIST/SEMATECH e-Handbook of Statistical Methods: Poisson distribution, accessed 1 October 2026. ↩
-
Meta Business Help Centre: About cost per conversion lift, accessed 1 October 2026. ↩