All articles

Measurement & Attribution

Test vs. control: designing a clean in-store media experiment

Test vs. control: designing a clean in-store media experiment

A test-versus-control experiment splits comparable stores into two groups, runs the campaign in one, holds it back from the other, and reads the difference in sales between them as the campaign's effect. It's the cleanest measurement design available to in-store media, and physical retail happens to be unusually well suited to it, because stores are discrete units you can assign and compare.

Digital channels spent years trying to approximate this with holdout audiences. In a store network, you can just do it.

Why do you need a control group at all?

Because the world doesn't pause for your flight. During any six-week advertising campaign, seasons shift, paydays land, weather turns, and rival brands run their own promotions. A before-and-after comparison in test stores alone tangles all of that with your media.

Control stores untangle it. They live through the same weeks and the same weather, so whatever moved sales everywhere shows up in both groups. The difference between the groups is what's left over, and that residue is your best estimate of what the campaign did.

How do you pick matched stores?

The goal is two groups that would have behaved the same without the campaign. Match on the traits that drive sales: store volume, category mix, retail channel, and geography. A high-volume urban bodega shouldn't be controlled by a quiet suburban grocer.

Scale helps here. With 34,000+ independently-owned stores across 8,500+ zip codes, the NRS network usually offers enough similar stores to build honest matches rather than settling for whatever's left. Because campaigns on NRS Digital Media activate by geography down to zip-code level and by retail channel, the assignment itself is a targeting decision, not a data engineering project.

One caution: don't let the test group be self-selected winners. If the campaign runs in your strongest stores and the controls are everyone else, the comparison is rigged before it starts.

What should stay constant during the test?

Everything you can hold. Price should match across groups. Trade promotions, display changes, and new distribution should either pause or apply equally to both. If a promotion hits only the test stores mid-flight, your experiment quietly becomes a promotion study with screens in the background.

You can't control everything, and that's fine. The rule is to log what you couldn't control. A one-page record of mid-flight changes lets you interpret an odd result later instead of arguing about half-remembered events.

How long should the test run, and what counts as an answer?

Run long enough to cover several full purchase cycles for the category. Products bought daily reveal themselves faster than products bought monthly. Resist reading results mid-flight and stopping when the numbers look good; a test that ends the moment it flatters you isn't a test.

Keep the flight boundaries aligned with whole weeks, too. Neighborhood store categories breathe on a weekly rhythm of paydays and weekend trips, and a window that cuts a week in half makes every baseline comparison messier than it needs to be.

At the end, compare each group's sales against its own baseline, then compare the two changes. Read the advertised SKUs specifically, and check whether the pattern holds across regions rather than resting on one lucky cluster. The transaction records that make this possible flow from the same point-of-sale platform the screens run on, so exposure and outcome come from one system.

Then, if the result is promising, do the most underrated thing in measurement: run it again.

Frequently asked questions

How many stores do I need for a valid test?

Enough that a few unusual stores can't dominate the read, and the right count depends on how variable the category is. The practical approach is to work with your media partner to size groups against the category's noise, then favor more stores over fewer when in doubt.

Can the control stores see any version of the campaign?

No. The control group's job is to show what would have happened with no exposure, so it needs genuinely dark screens for your brand during the flight. Since NRS is the exclusive media supplier in its locations, holdout stores stay clean rather than being reached by another provider.

What if my test result comes back flat?

Learn from it before rerunning it. Check whether the creative was suited to the screen, whether the flight covered enough purchase cycles, and whether mid-flight changes muddied the groups. A flat, clean test is real information about that campaign, not proof the channel can't work.