Ali Hosseini / data & analytics
Case study  ·  9 min read, 13 with the method sections open
E-commerce & retail / paid campaign optimisation

The best-performing campaigns were the ones doing the least work.

How a retailer stopped ranking campaigns by ROAS, and what it cost them to have been doing it.

Client anonymised. Spend and revenue indexed to a baseline; ratios and distributions unchanged. Analysis run September 2026.
revenue returned (indexed) 0 25 50 75 100 spend (indexed to the largest campaign) brand search spending here returns 0.08 on the next unit retargeting shopping prospecting still climbing

Scroll the chart sideways to see the full spend range.

Each dot is where the campaign was actually spending. Brand search and retargeting had the highest reported ROAS on the platform dashboard, and both sit on the flat part of their curve — the place where the next unit of budget buys almost nothing.
18moof spend and order history modelled
8campaigns given individual response curves, replacing one blended ROAS ranking
3campaigns found past the point of diminishing returns
22%of paid spend was estimated non-incremental1
On this page
  1. In one minute
  2. Marginal ROAS
  3. The problem
  4. What we had
  5. Approach
  6. Findings
  7. Why curves, not rankings
  8. What was delivered
  9. Recommendations
  10. Limitations
  11. How it was built
  12. The repository
  13. Running it yourself
01

In one minute

Budget was being allocated by last-click ROAS, which quietly funnelled money toward campaigns that harvest demand rather than create it. Fitting a response curve per campaign, and testing two of them with a geo holdout, changed where roughly a fifth of the paid budget goes.

02

What is marginal ROAS?

Marginal ROAS is what the next unit of budget returns. Average ROAS is what the dashboard reports. Average ROAS divides a campaign's total revenue by its total spend. Marginal ROAS is the slope of that campaign's response curve at the spend level it is currently running at.

The two are equal only if revenue rises in a straight line from zero as spend increases — which no advertising channel does. Every channel saturates, so average return stays high long after marginal return has collapsed. That gap is the entire subject of this case study: in this account, the campaign with the highest average ROAS returned an estimated 0.08 on its next unit of spend.

average = total revenue ÷ total spend  ·  marginal = dr/ds at current spend
Saturation curve

How revenue responds as spend rises. It bends because the pool of responsive customers is finite. Fitted here with a Hill function, one curve per campaign.

Geo holdout

Switching a campaign off across a matched subset of regions while the rest runs as control. Converts a modelled estimate into a measured one.

Incrementality

The share of conversions that would not have happened without the advertising. Platform-reported conversions are an upper bound, not a measurement.

Creative half-life

Days until a creative's click-through rate falls to half its opening rate. Measured at about 11 days across this account.

03

Why did ranking campaigns by ROAS break the budget?

Paid media has a structural trap that most reporting walks straight into. A campaign that shows an ad to someone already heading for checkout will report an excellent return, because the sale happens and the ad gets the credit. A campaign that introduces the brand to someone who buys six weeks later reports a poor one. Rank campaigns by return and you systematically move budget from the second kind to the first.

The retailer allocated budget monthly, by last-click ROAS, across eight campaigns. Top performers got more. It had been running that way for two years, and it had worked in the sense that reported ROAS improved every quarter — while new-customer orders had been flat for five.

Nobody had asked the question that breaks the loop: not which campaign returns the most, but what would the next unit of budget actually buy, in each one. Those are different questions and they have different answers.

Every campaign is above target. So why isn't the business growing?

Head of performance marketing, paraphrased from the kickoff session
04

What we had to work with

Eighteen months of daily campaign spend and platform-reported conversions, joined to order-level revenue from the store's own database, plus session data. Long enough to see spend vary meaningfully within each campaign, which is what a response curve needs — a campaign held at a flat budget all year teaches you nothing.

Three gaps shaped what the work could conclude:

  • One platform changed its attribution model mid-periodSix months of that campaign's reported conversions aren't comparable to the rest. Its curve is fitted on the later window only, with correspondingly wider uncertainty.
  • No view-through dataUpper-funnel video can only be credited where it produced a click or a measurable session. Its true contribution is understated here, and the recommendation is written to reflect that.
  • New versus returning wasn't resolvable at ad levelOrder-level data knows which customers were new. It can't reliably say which campaign reached them. That's why the acquisition question was answered with a holdout rather than a model.

The third gap is why this engagement included a live test rather than ending at analysis.

05

Questions, and how they were answered

  1. If we added one more unit of budget tomorrow, which campaign should get it — and which one is already full?
  2. How much of what our best-performing campaigns report would have happened without them?
  3. How long does a creative keep working before it stops paying for itself?
  4. What would the same total budget return if it were allocated differently?
Joinspend to real orders Fit response curvesone per campaign Derive marginal returnslope at current spend Testgeo holdout, 4 weeks Reallocateequalise marginal return
The highlighted stages are where the judgment lives. The reasoning is set out after the findings.

In plain terms: instead of asking what each campaign returned on average, the analysis fits the shape of each campaign's response to money — how revenue grows as spend grows — and reads the slope at the point where the campaign is currently sitting. That slope is what the next unit of budget buys, and it is usually nothing like the average. Where the curve made a strong claim, a four-week geo holdout was run to check it against reality rather than against the model.

06

What we found

The two highest-ROAS campaigns had the lowest marginal return of the eight

Brand search reported the best return in the account and returned an estimated 0.08 on the next unit of spend. Retargeting was second on both counts. Prospecting, which had been cut twice in the previous year for underperforming, had the highest marginal return of any campaign in the account.

Average return versus marginal return, by campaign
Both indexed to the account average; marginal is the slope of the fitted curve at current spend
1.5× 1.0× 0.5× 0.25× 0 brand retgt email px shopping pmax prospect video social campaigns, sorted by reported average return
average return — what the dashboard reports marginal return — what the next unit of budget buys
So what. The metric the budget process ran on was almost inversely related to the metric it should have run on. Every monthly reallocation for two years had been moving money in the wrong direction, with the reporting confirming it was working.

Cutting brand search in a third of the country cost 4% of its conversions

Rather than trust the curve alone, brand search was switched off across a matched third of postcodes for four weeks. Conversions in the test region fell briefly, recovered within about a fortnight, and settled roughly 4% below control — against a curve prediction of 6% and a platform-reported claim that the campaign was responsible for all of them.

Geo holdout: conversions in test versus control regions
Indexed to the pre-period; brand search switched off in the test region at week 8
1.2× 1.0× 0.8× 0.6× brand search off in test region wk 0 wk 8 wk 16 weeks
control regions — brand search left running test regions — brand search switched off at week 8
So what. The campaign was being credited with roughly twenty-five times the conversions it was actually causing. This is also the finding that made the rest credible internally: it was a real experiment on the client's own traffic, not a model output, and it moved the conversation from "your numbers disagree with the platform" to "the platform is wrong".

Creatives lost half their click-through within about eleven days

Tracking click-through by days since launch across the account's creative library, the decay is fast and consistent — a half-life around eleven days, with almost no recovery after. The refresh cadence at the time was monthly, meaning most creatives spent the back half of their life delivering at a third of their opening rate.

Click-through by creative age
Four representative creatives and the account median, indexed to day one
100% 75% 50% half-life: 11 days refresh happened here 0 7 14 21 28 days since the creative went live
account median individual creatives
So what. Half of the delivered impressions in a typical month came from creatives already past their half-life. Moving the refresh cadence to fortnightly is a production question, not a media-buying one, and it was the cheapest of the four recommendations to act on.

The same budget, reallocated, models out at 12% more revenue

Equalising marginal return across the eight campaigns — moving budget out of the flat parts of curves and into the steep ones — produces about 12% more revenue at identical total spend. Around 22% of the budget changes campaign under that allocation.

The 12% is a model output, and it should be read as a direction and a rough magnitude rather than a forecast. The holdout above is the only number on this page that was measured rather than modelled.

So what. The recommendation isn't "spend more". It's that the current allocation is leaving money on the table at the budget the business already approved, which is a materially easier conversation to have with finance.
07

Why curves, not rankings

Two decisions determined every number above. Both are open to inspection.

Why a response curve rather than a ROAS league table The simpler option, and why it produces the wrong answer every time

A ROAS ranking treats each campaign as a single number: total revenue over total spend. It's easy to compute, easy to present, and it answers a question nobody is actually asking. When you allocate budget you are never choosing between whole campaigns. You are choosing where the next unit goes.

Average and marginal return are the same number only when the relationship between spend and revenue is a straight line through the origin. It never is. Every paid channel saturates: there is a finite pool of people who will respond, and once you are reaching most of them, additional money buys duplicates and lower-intent impressions. The curve bends, and average return stays high long after marginal return has collapsed — which is exactly what makes a league table dangerous. It keeps rewarding the campaign whose curve flattened first.

Fitting a saturation curve per campaign and reading its slope at current spend answers the real question directly. Where campaigns sit on their curves also explains the pattern the client couldn't account for: reported ROAS rising every quarter while new customers stayed flat. The allocation process was optimising a number that improves as you approach saturation.

ROAS is still reported. It became a monitoring metric rather than an allocation input.

Why the analysis didn't stop at the model Where the curve was making its strongest claim, it was tested rather than trusted

Curves fitted to observational data carry a real risk: spend and demand move together. Budgets go up in peak season, and so do sales. A naive fit reads that shared seasonality as the campaign working, and the more a campaign's budget follows demand, the more overstated its curve becomes. Brand search is the worst case, because its spend rises precisely when people are already searching for the brand.

Bootstrapping the fit gives you the uncertainty in the curve, and the band widens sharply beyond the range of spend actually observed. That's useful and it is not sufficient — a confidently-fitted curve can still be confidently wrong if the confound is in the data itself.

So the campaign making the strongest claim was tested directly. A matched third of postcodes had brand search switched off for four weeks, with the rest as control. That converts the largest single line of spend from a modelled estimate into a measured one.

The result came in below the curve's prediction — 4% against 6% — which is the useful direction to be wrong in. It tells you the model is slightly optimistic about that campaign, and it sets a correction factor for the ones that couldn't be tested.

Fit uncertainty: shopping campaign
Bootstrap resampling, 1,000 draws; band widens beyond observed spend
observed spend range extrapolated 0 50 100 indexed spend
The band is why the reallocation recommends moving budget in steps rather than jumping straight to the modelled optimum.

Where it landed. Eight fitted curves, one measured holdout, and a reallocation applied in three monthly steps rather than one.

08

What the client actually received

  • Budget allocation modelTakes next month's total budget and returns a per-campaign split that equalises marginal return, with a cap on how far any campaign can move in one step. Used monthly by performance marketing.
  • Saturation monitorWhere each campaign currently sits on its curve, refreshed weekly. Flags a campaign when it crosses into the flat region.
  • Creative refresh calendarDays-since-launch tracking against the measured half-life, feeding the production schedule rather than the media plan.
  • Holdout playbookHow to run the next geo test, including region matching, minimum duration and the analysis script. So the second test doesn't need an external analyst.
Budget allocation monitor October plan Budget reallocated22% Modelled uplift+12% Campaigns past saturation3 of 8 Current split vs recommended brand retargeting prospecting Position on the response curve brand: saturated prospecting Creatives past their half-life shopping — 6 of 9 creativesavg age 19drefresh now prospecting — 2 of 7 creativesavg age 8don schedule social — 5 of 6 creativesavg age 24drefresh now
The top-left panel is the one that gets opened. It answers the only question the monthly budget meeting actually has: what should the split be, and how far is that from what we're doing now.
09

Recommendations

RecommendationExpected effectHow confident, and why
Cut brand search spend by roughly two thirdsFrees ~14% of total budget at a measured cost of 4% of that campaign's conversionsHighmeasured in a live geo holdout
Allocate by marginal return, not ROAS, in three monthly steps~12% more revenue at the same total spendMediummodelled; step-wise so it self-corrects
Move creative refresh from monthly to fortnightlyRemoves the tail half of the decay curve from deliveryHighdecay measured across the full library
Retest retargeting with the same holdout designConverts the second-largest modelled claim into a measured oneHighthe design is proven; only the campaign changes
Report ROAS as a monitoring metric, not an allocation inputBreaks the feedback loop that caused thisMediumprocess change; depends on adoption
10

What this study can't tell you

The 22% and the 12% are modelled, not measured. They come from fitted response curves, and a curve fitted to observational spend can mistake shared seasonality for campaign effect. One campaign — brand search, the largest single claim — was tested with a geo holdout and came in below its curve's prediction, which suggests the model is mildly optimistic across the board. Treat the reallocation figure as a direction and a rough magnitude, not a forecast. The stepped rollout exists so each month's result corrects the next month's allocation.

Upper-funnel video is understated throughout, because no view-through data was available and it can only be credited where it produced a click or a session. One campaign's curve is fitted on twelve months rather than eighteen, following a mid-period attribution change on that platform. Neither issue affects the finding about brand search, which was measured directly.

The creative decay figure is an account-level median. Individual creatives varied between roughly eight and fifteen days, and the analysis doesn't explain why — that would need creative-level tagging the account didn't have.

What I'd change: I would have started the geo holdout in week one rather than week three. It was the single most persuasive output of the engagement and it was also the long pole — four weeks of runtime meant the final recommendation landed a fortnight after the analysis was otherwise finished. Any future version of this work starts the test before the modelling, not after it.

Technical appendix Everything above stands on its own. What follows is for anyone who wants to check how it was built.
11

How it was built

Daily spend from three ad platforms is joined to order-level revenue from the store database, aggregated to campaign-day, and fitted with a saturation curve per campaign. Bootstrap resampling gives the uncertainty band. The whole fit runs in under two minutes; the slow part of the engagement was the four-week holdout, not the compute.

Data — BigQuery, Python 3.11, pandas
Analysis — SciPy (curve_fit), NumPy, statsmodels
Methods — Hill saturation function, bootstrap resampling, difference-in-differences
Serving — Looker Studio, weekly scheduled refresh
Reproducibility — Jupyter, parameterised with papermill

View the implementation

Ad platform APIs3 sources, daily Store databaseorder-level revenue Campaign-day tablespend + real revenue Curve fit + bootstrapone per campaign Marginal returnslope at current spend Allocation modelstepped, capped Looker Studio Geo holdoutmanual, quarterly
The holdout sits outside the automated path on purpose. It's a quarterly decision, not a scheduled job, and treating it as one is how these tests stop happening.
See the core logic — fitting the curve and reading the marginal return
import numpy as np
from scipy.optimize import curve_fit


def response(spend, vmax, half, shape):
    """
    Hill saturation:  r(s) = vmax · s^shape / (half^shape + s^shape)

    'half' is the spend at which the campaign returns half its
    ceiling. 'shape' controls how sharply the curve bends — above 1
    there is an S-shaped ramp before saturation, which is what
    upper-funnel campaigns actually do and what a linear ROAS
    ranking cannot represent at all.
    """
    return vmax * spend**shape / (half**shape + spend**shape)


def marginal_return(spend, params, h=1e-3):
    """What the NEXT unit of budget buys — the slope, not the average.

    This is the whole argument of the engagement in one line: the
    budget decision is a question about the derivative, and ROAS
    reports the secant from the origin.
    """
    return (response(spend + h, *params) - response(spend - h, *params)) / (2 * h)


def fit_campaign(daily, n_boot=1000, seed=0):
    """Fit one campaign, with a bootstrap band over the fitted curve."""
    s, r = daily["spend"].values, daily["revenue"].values
    p0 = [r.max() * 1.2, np.median(s), 1.0]
    params, _ = curve_fit(response, s, r, p0=p0, maxfev=20000,
                          bounds=([0, 1e-6, 0.5], [np.inf, np.inf, 3.0]))

    rng = np.random.default_rng(seed)
    draws = []
    for _ in range(n_boot):
        idx = rng.integers(0, len(s), len(s))
        try:
            d, _ = curve_fit(response, s[idx], r[idx], p0=params, maxfev=20000)
            draws.append(d)
        except RuntimeError:
            continue

    # The band widens fast outside the observed spend range. That is a
    # feature: it is why the reallocation moves in steps rather than
    # jumping to the modelled optimum, which sits in extrapolated space.
    return params, np.array(draws), (s.min(), s.max())


current = {c: marginal_return(spend_now[c], fits[c]) for c in campaigns}

# {'brand': 0.08, 'retargeting': 0.15, 'shopping': 0.68, 'prospecting': 0.90}
# Reported ROAS, same order:  1.19, 1.28, 0.92, 0.83
12

Check the work

campaign-response-curves

Saturation curve fitting, marginal return, and budget allocation for multi-campaign paid media, plus the geo holdout analysis.

  • Hill curve fitting with bootstrap confidence bands
  • Marginal return and the stepped allocation optimiser
  • Geo holdout: region matching, power calculation, and the difference-in-differences analysis
  • Synthetic data generator — 18 months, 8 campaigns, with seasonality deliberately confounded into spend so the failure mode is visible

No client data is included. Running the generator reproduces the analysis end to end and the shape of every chart on this page — the values differ, because the client's data isn't in the repository.

View the repository
13

If you want to run this yourself

Four things I'd tell anyone attempting this analysis on their own account, including the parts that catch people out. The repository above has the code; this is the judgment that doesn't fit in it.

How do you know if a campaign is past the point of diminishing returns?

Fit a saturation curve to daily spend and revenue, then read the slope where the campaign currently sits. A near-flat slope means additional budget buys very little.

The precondition catches most people out: spend has to have varied meaningfully inside the observation window. A campaign held at a flat budget all year produces no usable curve, which is why this needed eighteen months of history rather than a quarter.

Is branded search actually incremental?

Usually far less than platform reporting claims, because branded search spend rises exactly when people are already searching for the brand. The confound is built into the data, so no amount of modelling removes it.

Treat platform-reported conversions for branded search as an upper bound. The only way to get the real figure for your account is to switch it off somewhere and measure.

How long should a geo holdout run?

Longer than feels necessary. Split regions into matched test and control groups, switch the campaign off in the test group only, and index conversions to a pre-period baseline before comparing.

Four weeks was the right length here because the initial drop recovered substantially within about a fortnight. A two-week test would have measured the shock rather than the settled effect, and reported a far larger number than the true one.

How often should creative be refreshed?

Measure your own half-life rather than adopting someone else's cadence. Track click-through by days since launch across your library and find the point where it falls to half its opening rate.

If that half-life is shorter than your refresh cycle, a large share of your delivered impressions is coming from creative already past its useful life — which is a production scheduling problem, not a media buying one.

If your best-performing campaign is your branded search, you're probably paying for traffic you already had.

Most paid accounts allocate budget with a metric that rewards demand harvesting and penalises demand creation. The cost isn't a wrong number on a dashboard. It's that the number improves every quarter while the business stops growing.

I work with e-commerce and subscription businesses on incrementality testing, budget allocation, and the measurement infrastructure underneath both.

Send me your campaign list and last quarter's spend split, and in half an hour I'll tell you which campaigns are likely past saturation and which single holdout test would be worth running first. No deck, no proposal. If your allocation already looks sound, I'll say so.