In one minute
Budget was being allocated by last-click ROAS, which quietly funnelled money toward campaigns that harvest demand rather than create it. Fitting a response curve per campaign, and testing two of them with a geo holdout, changed where roughly a fifth of the paid budget goes.
What is marginal ROAS?
Marginal ROAS is what the next unit of budget returns. Average ROAS is what the dashboard reports. Average ROAS divides a campaign's total revenue by its total spend. Marginal ROAS is the slope of that campaign's response curve at the spend level it is currently running at.
The two are equal only if revenue rises in a straight line from zero as spend increases — which no advertising channel does. Every channel saturates, so average return stays high long after marginal return has collapsed. That gap is the entire subject of this case study: in this account, the campaign with the highest average ROAS returned an estimated 0.08 on its next unit of spend.
average = total revenue ÷ total spend · marginal = dr/ds at current spendHow revenue responds as spend rises. It bends because the pool of responsive customers is finite. Fitted here with a Hill function, one curve per campaign.
Switching a campaign off across a matched subset of regions while the rest runs as control. Converts a modelled estimate into a measured one.
The share of conversions that would not have happened without the advertising. Platform-reported conversions are an upper bound, not a measurement.
Days until a creative's click-through rate falls to half its opening rate. Measured at about 11 days across this account.
Why did ranking campaigns by ROAS break the budget?
Paid media has a structural trap that most reporting walks straight into. A campaign that shows an ad to someone already heading for checkout will report an excellent return, because the sale happens and the ad gets the credit. A campaign that introduces the brand to someone who buys six weeks later reports a poor one. Rank campaigns by return and you systematically move budget from the second kind to the first.
The retailer allocated budget monthly, by last-click ROAS, across eight campaigns. Top performers got more. It had been running that way for two years, and it had worked in the sense that reported ROAS improved every quarter — while new-customer orders had been flat for five.
Nobody had asked the question that breaks the loop: not which campaign returns the most, but what would the next unit of budget actually buy, in each one. Those are different questions and they have different answers.
Every campaign is above target. So why isn't the business growing?
Head of performance marketing, paraphrased from the kickoff session
What we had to work with
Eighteen months of daily campaign spend and platform-reported conversions, joined to order-level revenue from the store's own database, plus session data. Long enough to see spend vary meaningfully within each campaign, which is what a response curve needs — a campaign held at a flat budget all year teaches you nothing.
Three gaps shaped what the work could conclude:
- One platform changed its attribution model mid-periodSix months of that campaign's reported conversions aren't comparable to the rest. Its curve is fitted on the later window only, with correspondingly wider uncertainty.
- No view-through dataUpper-funnel video can only be credited where it produced a click or a measurable session. Its true contribution is understated here, and the recommendation is written to reflect that.
- New versus returning wasn't resolvable at ad levelOrder-level data knows which customers were new. It can't reliably say which campaign reached them. That's why the acquisition question was answered with a holdout rather than a model.
The third gap is why this engagement included a live test rather than ending at analysis.
Questions, and how they were answered
- If we added one more unit of budget tomorrow, which campaign should get it — and which one is already full?
- How much of what our best-performing campaigns report would have happened without them?
- How long does a creative keep working before it stops paying for itself?
- What would the same total budget return if it were allocated differently?
In plain terms: instead of asking what each campaign returned on average, the analysis fits the shape of each campaign's response to money — how revenue grows as spend grows — and reads the slope at the point where the campaign is currently sitting. That slope is what the next unit of budget buys, and it is usually nothing like the average. Where the curve made a strong claim, a four-week geo holdout was run to check it against reality rather than against the model.
What we found
The two highest-ROAS campaigns had the lowest marginal return of the eight
Brand search reported the best return in the account and returned an estimated 0.08 on the next unit of spend. Retargeting was second on both counts. Prospecting, which had been cut twice in the previous year for underperforming, had the highest marginal return of any campaign in the account.
Cutting brand search in a third of the country cost 4% of its conversions
Rather than trust the curve alone, brand search was switched off across a matched third of postcodes for four weeks. Conversions in the test region fell briefly, recovered within about a fortnight, and settled roughly 4% below control — against a curve prediction of 6% and a platform-reported claim that the campaign was responsible for all of them.
Creatives lost half their click-through within about eleven days
Tracking click-through by days since launch across the account's creative library, the decay is fast and consistent — a half-life around eleven days, with almost no recovery after. The refresh cadence at the time was monthly, meaning most creatives spent the back half of their life delivering at a third of their opening rate.
The same budget, reallocated, models out at 12% more revenue
Equalising marginal return across the eight campaigns — moving budget out of the flat parts of curves and into the steep ones — produces about 12% more revenue at identical total spend. Around 22% of the budget changes campaign under that allocation.
The 12% is a model output, and it should be read as a direction and a rough magnitude rather than a forecast. The holdout above is the only number on this page that was measured rather than modelled.
Why curves, not rankings
Two decisions determined every number above. Both are open to inspection.
Why a response curve rather than a ROAS league table The simpler option, and why it produces the wrong answer every time
A ROAS ranking treats each campaign as a single number: total revenue over total spend. It's easy to compute, easy to present, and it answers a question nobody is actually asking. When you allocate budget you are never choosing between whole campaigns. You are choosing where the next unit goes.
Average and marginal return are the same number only when the relationship between spend and revenue is a straight line through the origin. It never is. Every paid channel saturates: there is a finite pool of people who will respond, and once you are reaching most of them, additional money buys duplicates and lower-intent impressions. The curve bends, and average return stays high long after marginal return has collapsed — which is exactly what makes a league table dangerous. It keeps rewarding the campaign whose curve flattened first.
Fitting a saturation curve per campaign and reading its slope at current spend answers the real question directly. Where campaigns sit on their curves also explains the pattern the client couldn't account for: reported ROAS rising every quarter while new customers stayed flat. The allocation process was optimising a number that improves as you approach saturation.
ROAS is still reported. It became a monitoring metric rather than an allocation input.
Why the analysis didn't stop at the model Where the curve was making its strongest claim, it was tested rather than trusted
Curves fitted to observational data carry a real risk: spend and demand move together. Budgets go up in peak season, and so do sales. A naive fit reads that shared seasonality as the campaign working, and the more a campaign's budget follows demand, the more overstated its curve becomes. Brand search is the worst case, because its spend rises precisely when people are already searching for the brand.
Bootstrapping the fit gives you the uncertainty in the curve, and the band widens sharply beyond the range of spend actually observed. That's useful and it is not sufficient — a confidently-fitted curve can still be confidently wrong if the confound is in the data itself.
So the campaign making the strongest claim was tested directly. A matched third of postcodes had brand search switched off for four weeks, with the rest as control. That converts the largest single line of spend from a modelled estimate into a measured one.
The result came in below the curve's prediction — 4% against 6% — which is the useful direction to be wrong in. It tells you the model is slightly optimistic about that campaign, and it sets a correction factor for the ones that couldn't be tested.
Where it landed. Eight fitted curves, one measured holdout, and a reallocation applied in three monthly steps rather than one.
What the client actually received
- Budget allocation modelTakes next month's total budget and returns a per-campaign split that equalises marginal return, with a cap on how far any campaign can move in one step. Used monthly by performance marketing.
- Saturation monitorWhere each campaign currently sits on its curve, refreshed weekly. Flags a campaign when it crosses into the flat region.
- Creative refresh calendarDays-since-launch tracking against the measured half-life, feeding the production schedule rather than the media plan.
- Holdout playbookHow to run the next geo test, including region matching, minimum duration and the analysis script. So the second test doesn't need an external analyst.
Recommendations
| Recommendation | Expected effect | How confident, and why |
|---|---|---|
| Cut brand search spend by roughly two thirds | Frees ~14% of total budget at a measured cost of 4% of that campaign's conversions | Highmeasured in a live geo holdout |
| Allocate by marginal return, not ROAS, in three monthly steps | ~12% more revenue at the same total spend | Mediummodelled; step-wise so it self-corrects |
| Move creative refresh from monthly to fortnightly | Removes the tail half of the decay curve from delivery | Highdecay measured across the full library |
| Retest retargeting with the same holdout design | Converts the second-largest modelled claim into a measured one | Highthe design is proven; only the campaign changes |
| Report ROAS as a monitoring metric, not an allocation input | Breaks the feedback loop that caused this | Mediumprocess change; depends on adoption |
What this study can't tell you
The 22% and the 12% are modelled, not measured. They come from fitted response curves, and a curve fitted to observational spend can mistake shared seasonality for campaign effect. One campaign — brand search, the largest single claim — was tested with a geo holdout and came in below its curve's prediction, which suggests the model is mildly optimistic across the board. Treat the reallocation figure as a direction and a rough magnitude, not a forecast. The stepped rollout exists so each month's result corrects the next month's allocation.
Upper-funnel video is understated throughout, because no view-through data was available and it can only be credited where it produced a click or a session. One campaign's curve is fitted on twelve months rather than eighteen, following a mid-period attribution change on that platform. Neither issue affects the finding about brand search, which was measured directly.
The creative decay figure is an account-level median. Individual creatives varied between roughly eight and fifteen days, and the analysis doesn't explain why — that would need creative-level tagging the account didn't have.
What I'd change: I would have started the geo holdout in week one rather than week three. It was the single most persuasive output of the engagement and it was also the long pole — four weeks of runtime meant the final recommendation landed a fortnight after the analysis was otherwise finished. Any future version of this work starts the test before the modelling, not after it.
How it was built
Daily spend from three ad platforms is joined to order-level revenue from the store database, aggregated to campaign-day, and fitted with a saturation curve per campaign. Bootstrap resampling gives the uncertainty band. The whole fit runs in under two minutes; the slow part of the engagement was the four-week holdout, not the compute.
Data — BigQuery, Python 3.11, pandas
Analysis — SciPy (curve_fit), NumPy, statsmodels
Methods — Hill saturation function, bootstrap resampling, difference-in-differences
Serving — Looker Studio, weekly scheduled refresh
Reproducibility — Jupyter, parameterised with papermill
See the core logic — fitting the curve and reading the marginal return
import numpy as np from scipy.optimize import curve_fit def response(spend, vmax, half, shape): """ Hill saturation: r(s) = vmax · s^shape / (half^shape + s^shape) 'half' is the spend at which the campaign returns half its ceiling. 'shape' controls how sharply the curve bends — above 1 there is an S-shaped ramp before saturation, which is what upper-funnel campaigns actually do and what a linear ROAS ranking cannot represent at all. """ return vmax * spend**shape / (half**shape + spend**shape) def marginal_return(spend, params, h=1e-3): """What the NEXT unit of budget buys — the slope, not the average. This is the whole argument of the engagement in one line: the budget decision is a question about the derivative, and ROAS reports the secant from the origin. """ return (response(spend + h, *params) - response(spend - h, *params)) / (2 * h) def fit_campaign(daily, n_boot=1000, seed=0): """Fit one campaign, with a bootstrap band over the fitted curve.""" s, r = daily["spend"].values, daily["revenue"].values p0 = [r.max() * 1.2, np.median(s), 1.0] params, _ = curve_fit(response, s, r, p0=p0, maxfev=20000, bounds=([0, 1e-6, 0.5], [np.inf, np.inf, 3.0])) rng = np.random.default_rng(seed) draws = [] for _ in range(n_boot): idx = rng.integers(0, len(s), len(s)) try: d, _ = curve_fit(response, s[idx], r[idx], p0=params, maxfev=20000) draws.append(d) except RuntimeError: continue # The band widens fast outside the observed spend range. That is a # feature: it is why the reallocation moves in steps rather than # jumping to the modelled optimum, which sits in extrapolated space. return params, np.array(draws), (s.min(), s.max()) current = {c: marginal_return(spend_now[c], fits[c]) for c in campaigns} # {'brand': 0.08, 'retargeting': 0.15, 'shopping': 0.68, 'prospecting': 0.90} # Reported ROAS, same order: 1.19, 1.28, 0.92, 0.83
Check the work
campaign-response-curves
Saturation curve fitting, marginal return, and budget allocation for multi-campaign paid media, plus the geo holdout analysis.
- Hill curve fitting with bootstrap confidence bands
- Marginal return and the stepped allocation optimiser
- Geo holdout: region matching, power calculation, and the difference-in-differences analysis
- Synthetic data generator — 18 months, 8 campaigns, with seasonality deliberately confounded into spend so the failure mode is visible
No client data is included. Running the generator reproduces the analysis end to end and the shape of every chart on this page — the values differ, because the client's data isn't in the repository.
View the repositoryIf you want to run this yourself
Four things I'd tell anyone attempting this analysis on their own account, including the parts that catch people out. The repository above has the code; this is the judgment that doesn't fit in it.
How do you know if a campaign is past the point of diminishing returns?
Fit a saturation curve to daily spend and revenue, then read the slope where the campaign currently sits. A near-flat slope means additional budget buys very little.
The precondition catches most people out: spend has to have varied meaningfully inside the observation window. A campaign held at a flat budget all year produces no usable curve, which is why this needed eighteen months of history rather than a quarter.
Is branded search actually incremental?
Usually far less than platform reporting claims, because branded search spend rises exactly when people are already searching for the brand. The confound is built into the data, so no amount of modelling removes it.
Treat platform-reported conversions for branded search as an upper bound. The only way to get the real figure for your account is to switch it off somewhere and measure.
How long should a geo holdout run?
Longer than feels necessary. Split regions into matched test and control groups, switch the campaign off in the test group only, and index conversions to a pre-period baseline before comparing.
Four weeks was the right length here because the initial drop recovered substantially within about a fortnight. A two-week test would have measured the shock rather than the settled effect, and reported a far larger number than the true one.
How often should creative be refreshed?
Measure your own half-life rather than adopting someone else's cadence. Track click-through by days since launch across your library and find the point where it falls to half its opening rate.
If that half-life is shorter than your refresh cycle, a large share of your delivered impressions is coming from creative already past its useful life — which is a production scheduling problem, not a media buying one.