Reading campaign results without fooling yourself
Reading campaign results honestly comes down to three habits: compare against a real baseline, account for everything else that changed during the flight, and decide what success looks like before the numbers arrive. Skip any one of them and the data will happily tell you whatever you were hoping to hear. The easiest person for a results deck to fool is its author.
None of this requires advanced statistics. It requires discipline, mostly at the moments when discipline is least fun.
Why does every result need a baseline?
Because a number on its own means nothing. Sales rose in advertised stores; compared with what? A fair baseline gives the number its meaning, and the strongest one is a control group of matched stores that didn't run the campaign through the same weeks. Their trajectory shows what your test stores would likely have done anyway.
Weaker baselines invite self-flattery. A before-and-after read credits the campaign with every tailwind that happened to blow during the flight. If a control group wasn't possible, a year-ago comparison over the same weeks is the minimum honest substitute.
What else changed while your campaign ran?
More than you'd think, and the calendar is the chief suspect. Categories in neighborhood retail move with seasons, holidays, and paydays; a flight that spans the start of summer will look brilliant for cold drinks no matter what was on the screens. Channels like convenience stores have their own rhythms, and any read that ignores them is measuring the calendar, not the campaign.
The same goes for prices, pack changes, new distribution, and trade promotions. Keep a running log of these during the flight. When a result looks surprising in either direction, that log is usually where the explanation lives.
Are you cherry-picking without noticing?
Cherry-picking rarely feels dishonest from the inside. It feels like "focusing on what worked." Suppose a campaign ran in twelve regions and one shone, so the deck leads with that region. Twenty SKUs were advertised and three moved, so those three become the story.
Slice enough ways and something will always look good by chance. The honest structure is the opposite: report the overall result first, exactly as the pre-launch plan defined it, and then explore the slices as hypotheses for the next flight. A standout region is a lead to test, not a conclusion to bank.
What should be decided before launch?
The success criteria, in writing. Which SKUs count, which stores are in the test and control groups, what time window applies, and what size of difference you'd consider meaningful. This is the measurement version of aiming before firing, and it's what makes the eventual answer credible to a skeptic.
It also keeps vocabulary honest. If the plan said sales lift in advertised SKUs and the flight delivered a nice-looking engagement anecdote instead, the campaign missed, and the report should say so plainly. Since campaigns on NRS Digital Media can be structured by SKU, geography, and retail channel from the start, the test design can be built into the buy rather than reverse-engineered afterward.
When is a result real?
When it repeats. One flight producing one favorable read is a promising data point. The same design producing a similar read in a second flight, or holding across store types and regions, is a finding you can plan budgets on. Transaction-level reporting of the kind NRS Insights works with makes repetition cheap to check, because the outcome data keeps flowing after the campaign ends.
Patience here pays twice: you avoid scaling a fluke, and when you do scale, you can defend the decision with more than enthusiasm.
Frequently asked questions
What's the most common way teams fool themselves with results?
Comparing the flight only against the weeks before it. Seasonality, promotions, and price moves all ride along in that comparison, and the campaign absorbs credit for them. A matched control group, or at minimum a year-ago window, strips most of that borrowed credit away.
Is it wrong to highlight the best-performing stores?
It's fine to explore them and wrong to lead with them. The headline result should be the pre-defined comparison across the full test. Standout stores are hypotheses about where the campaign works best, worth testing deliberately in the next flight rather than presenting as proof.
How do I judge results with no industry benchmark to compare against?
Judge against your own plan and your own history. A pre-launch success threshold, a clean control comparison, and a repeated result across flights tell you more than any borrowed number, since published benchmarks rarely share your category, stores, creative, or measurement method.