Did the campaign actually work?
Two marketing campaigns at a leading Nigerian fintech, a rewards campaign and a TV campaign. Both "went up" on a before-and-after chart. The CMO asked me whether they had actually worked, which meant asking how much they went up compared with what would have happened anyway.
Situation
A rewards campaign and a TV campaign had both run, and both looked good on before-and-after charts. But the business was growing anyway, so those charts couldn't say what the campaigns had added.
Task
The CMO asked me to measure what each campaign actually did, and which part of behaviour it changed: how many people transacted, or how much each person did.
Action
Built a "no campaign" counterfactual with CausalImpact from customers outside the campaign, tested two control groups at two time grains, and used trend projections for TV, which reached everyone.
Result
The rewards campaign lifted daily active users 17–19% and volume 16–22%, but transactions per user didn't move. The next campaign was redesigned to target frequency.
The problem with before-and-after
Marketing teams usually judge a campaign by putting the weeks before launch next to the weeks after. If the line goes up, the campaign worked. But the business was growing anyway. Customers were onboarding every day, salaries land at month end, and some products grow on their own. A before-and-after chart credits all of that to the campaign.
What you actually want is the counterfactual: what the same customers would have done over the same days if the campaign had never run. You can't observe it, so you have to estimate it, and the estimate is only as good as what you build it from.
Charts use illustrative data rebuilt from the shape of the real results: values are indexed and lightly perturbed, and the real figures stay with the company. The last figure is a pure simulation and is labelled as such.
Story A: the rewards campaign
In October 2024 a set of personal banking customers was assigned to a rewards campaign covering three everyday products: airtime top-ups, interbank transfers (sending money to an account at another bank) and deposits. Everyone else was not assigned. That gave me a natural comparison group.
I used CausalImpact, Google's Bayesian structural time-series method. In plain terms it works like this:
- Learn the relationship before launch. In the pre-period, fit how the campaign group's daily series moves with the control group's series, plus trend and weekly seasonality.
- Project it forward. After launch, feed in the control series (which the campaign didn't touch) and predict what the campaign group "should" have done. That prediction is the counterfactual, with an uncertainty band.
- Measure the gap. Actual minus counterfactual, day by day and cumulatively. If the gap sits outside the band, the effect is unlikely to be noise.
How I built it
- The data. For each day from the start of August to late October 2024, and for each of the three products, I pulled the same three numbers for the campaign group and for the comparison group: active users, transaction volume, and transactions per user. I built a weekly version too, from June.
- The windows. August and September were the pre-period the model learned from; the campaign weeks in October were the post-period. The last week in the extract was only partly complete, so I cut it rather than let a half-week read as a collapse.
- The models. I used pycausalimpact, the Python port of Google's package, with the comparison group's series as the control. One model for each product and metric is nine models; I ran that set against two different comparison groups and at two time grains, 36 models in all, and read their summaries and plots side by side.
The figure below runs a simplified version of the same idea (ordinary least squares on the control series instead of the full Bayesian model) on series rebuilt from the real daily data. Switch to teaching mode to plant an effect of known size on the real control series and see when the method finds it and when it doesn't.
Daily series for customers assigned to the campaign and for everyone else, rebuilt from the real data, indexed (the unassigned group's pre-launch average = 100) and lightly perturbed. Note the weekly rhythm and the quiet Sundays. A regression fitted on the pre-period predicts the assigned group after launch (day 0) from the unassigned group; the shaded band is ±1.96 residual standard deviations. Bottom: cumulative extra activity over the post-period, which is short in this extract. In teaching mode the campaign group is rebuilt from the real control series with extra noise and a planted effect: try high noise with a short pre-period and watch the estimate drift. "Control + own trend" also lets the counterfactual carry on the assigned group's own pre-launch growth.
What the rewards campaign did
Against the counterfactual, the campaign group showed a 17–19% lift in daily active users across the three products, and a 16–22% lift in transaction volume. Both were well outside the uncertainty band. Airtime and deposits jumped in the first few days of the campaign and held that level. Interbank transfers eased off slightly once the campaign ended, after the short window shown here.
Look at the run-up in the figure, though. The assigned group was already growing faster than everyone else for months before launch. A counterfactual built only from the control series reads all of that extra growth after launch as campaign effect. Let the model carry on the assigned group's own trend and the lift shrinks to single digits, and for interbank volume it all but disappears. Frequency stays flat either way. I'd report both today; at the time I reported the first.
The third metric told a different story. Transactions per user didn't move. The model's estimated effect was within a few points of zero on every product.
Why that matters
The campaign brought more people in, not more activity per person. Volume rose only because more customers transacted. Each one transacted about as often as before. Looking back at the messaging, it rewarded taking part, not transacting more often. It never asked for frequency, so it didn't get any. That gave the next campaign a clear brief: target frequency explicitly and measure transactions per user as the primary outcome. The CMO took it on, and the next campaign was redesigned around frequency.
The choice of control changes the answer
I didn't stop at one model. I re-ran the analysis with two different control groups and at two time grains:
| Variant | What it compares | What happened |
|---|---|---|
| All-customer control, daily | Campaign group vs all unassigned customers | Users +17–19%, volume +16–22% across the three products. Frequency flat. |
| Personal-only control, daily | Campaign group vs unassigned personal banking customers only | Much less stable. One product's lift became negative and not significant, while others grew to several times the main estimate. |
| Weekly grain | Same designs, aggregated to weeks | Only two post-period weeks, so the intervals are wide. Against personal-only customers, the airtime user lift shrank to about 1.5% and was not distinguishable from zero. |
The direction held up across variants: users and volume up, frequency flat. The size didn't. A control group that looks "more comparable" on paper can track the treated series worse in the pre-period, and then the counterfactual goes wrong. The lesson I took away: the control is a modelling decision, not a detail. It has to be justified (how well does it track the treated series before launch?) and the sensitivity has to be reported, not hidden behind the one result that looks best.
Story B: the TV campaign
The TV campaign was harder. It reached everyone, so there was no unexposed group to compare against. I split customers into segments: active (transacting every month), new (first-ever transaction in the period) and returning (came back after a month or more away). I split each by whether they had come in through a referral. For each segment I took the weekly growth rate before the campaign, projected it forward as the "no TV" line, and compared actuals against it.
- Short first-week lift. New customers, referred or not, beat the projection in the first week of airing. Active customers were on or slightly above it.
- Then below projection. From the second week, new customers fell below the projected line and stayed there. Actives tracked it for a while, then slipped below in the last weeks.
- Returning customers were the bright spot. They had been declining steeply week on week before the campaign. During it, that decline slowed. It didn't reverse, but by the later weeks fewer people had drifted away than the trend predicted.
Weekly active customers in each segment, indexed so the baseline week is 100, rebuilt from the original report's weekly figures and lightly perturbed. The dashed line compounds the segment's pre-campaign weekly growth rate forward, as the original analysis did. New customers start above the line and fall away; actives track it and then slip below; returning customers, falling fast before the campaign, end above it.
My honest takeaway at the time was a modest impact. TV seemed to sustain visibility and slow the leak of lapsed customers rather than drive an immediate jump.
Looking back, the method deserves more scrutiny than the result. A constant compounding growth rate keeps accelerating in absolute terms, while real growth usually slows. So the projection overstates the counterfactual, and "actuals below projection" may partly just be the projection being wrong. The next figure shows how much the choice of baseline can swing the answer.
Weekly active users for a growing business. With the true effect at 0%, a before-and-after comparison still reports a large "lift" because the trend carries on. A compounded projection is right when growth is steady, but overstates the no-campaign line when growth is slowing, so a real effect can look like underperformance. The control-based counterfactual follows what the market actually did.
What I'd do differently
- Use a synthetic control for TV. With no unexposed group, I'd build one from weighted combinations of series the campaign plausibly didn't move. Examples are regions with little airtime, products not advertised, or a comparable market. That beats extrapolating a growth rate.
- Pre-register the control and the metric. Decide the control group, grain and primary metric before looking at post-launch data. Then report the alternatives as sensitivity checks.
- Check pre-period fit before reading the effect. If the counterfactual can't track the treated series before launch, its post-launch gap means little. Here the assigned group was out-growing the control before launch, which a control-only model quietly credits to the campaign.
- Hold out a randomised group. For incentive campaigns that are under our control, a small random holdout makes most of this modelling unnecessary. That's the direction the later experiment work took.
- Watch for novelty effects. First-week spikes that fade should be reported separately from the sustained effect. Averaging them together flatters short campaigns.