Causal Intelligence Adds a New Layer of Insight to Multichannel Advertising

By Matt Emans, CTO and Co-Founder at Newton Research

AI is making it easier for advertisers and agencies to tap into measurement techniques that were previously expensive, complex and time consuming, limiting their utility in timely optimization and scenario planning. In the past year, more advertisers have employed AI to help with measurement approaches including multi-touch attribution (MTA), media mix modeling (MMM) and incrementality. What used to cost millions and take months now fits into a normal budget and takes days or hours to complete.

But cheaper, faster and more accessible versions of the same methods still miss many of the signals and optimization opportunities across media. MMM answers a channel-level question: how much should go to CTV versus search versus linear, usually at a weekly grain and usually on a rebuild cycle measured in quarters. MTA answers a touchpoint question: of the impressions we can observe, which ones get credit — allocated algorithmically, not by counterfactual, and blind to whether the outcome would have happened anyway. Between the two sits the part of the plan planners actually spend their time on: the campaigns, audiences, creatives, dayparts and flighting decisions that make up a channel’s budget. MMM is too coarse to see them. MTA can see them but can’t tell you which ones caused anything.

That middle layer is what the evolution of causal modeling opens up. Causal discovery methods, Bayesian inference and continuous experimentation — now practical at scale because of AI, compute availability and deep learning techniques — make it possible to estimate the incremental effect of individual elements of a plan rather than assign credit after the fact. Instead of “shift 10% from display to CTV,” a brand can ask what happens if it moves 10% within a channel, across specific campaigns, creatives and audiences, and get an answer with a confidence interval attached.

That said, it’s not always the right answer. It requires historical data records along with key attributes and frequent observations of the KPIs that can be tested accurately using scientific methods.

How Causal Modeling Works

A causal model shouldn’t start from a blank page. Most large advertisers have already invested years in a media mix model, and that work carries something a finer-grained model can’t derive on its own: a view of the whole business, including channels with no impression-level record at all — linear, audio, out-of-home — plus baseline demand, seasonality, price and competitive pressure. A granular model that ignores all of that will systematically over-credit whatever it can see. So the MMM becomes the constraint rather than the competition. It sets the channel-level envelope, and the causal layer works inside it.

The second requirement is where a lot of current work goes wrong. Deep learning makes it remarkably easy to fit observed media data — plenty of degrees of freedom, plenty of history, and a large enough network will reproduce last year’s sales curve with impressive accuracy. 

But that’s a trap! Predictive fit tells you what tends to accompany what; it says nothing reliable about what happens when you make a change to a budget or tactic, and changing something is the only reason a planner opens the model. A model trained on the fact that spend rises ahead of holiday sales will cheerfully imply that spending more causes the holiday. The test that matters isn’t goodness of fit — it’s whether the model can answer “what happens if we move this budget” in a way that holds up when you actually move it, with honest uncertainty attached. That requires explicit causal structure and experimentation, and it means accepting a slightly worse fit in exchange for an answer you can act on.

What’s changed is that the machinery now runs at the scale media data actually arrives in. A national advertiser generates billions of records a year — impressions, bids, site events, transactions, plus everything moving in the market around them. Causal methods couldn’t touch that volume until recently: discovery algorithms scaled badly as candidate variables multiplied, and the Bayesian models that produce honest uncertainty took days to fit. GPU acceleration, modern inference and sequence architectures like the transformer have collapsed those runtimes and made it possible to learn from long event histories rather than weekly aggregates. That changes what’s askable — a model can search for structure across thousands of variables instead of a few dozen hand-curated ones, refit often enough to keep pace with a plan that changes weekly, and simulate a proposed set of changes forward rather than only explain the plan that already ran.

Once those pieces are in place, the model can do what an MMM can’t. Take a brand running dozens of campaigns on Amazon. The causal layer identifies pockets of over- and under-performance across specific campaigns, audiences and creatives and returns budget adjustments for a human-in-the-loop workflow — while still accounting for interaction with the CTV halo, the linear spend it can’t observe directly, and the overall budget trade-off. What comes back isn’t a credit allocation. It’s a what-if scenario at the grain where the decision actually gets made.

Working With the Data You Actually Have

The practical obstacle usually isn’t method, it’s resolution. Some media data is genuinely user-level — site behavior, retail media exposure, first-party CRM. Considerably more is aggregated by the time anyone can use it: platform reporting rolled up to campaign-by-day, walled-garden exports with no user key, panel-based TV data, clean-room outputs that are aggregate by design. Both standard responses throw away information. Aggregate everything to the lowest common denominator and you discard the granular signal you paid for. Insist on user-level modeling and you can only measure the addressable fraction of the plan, which biases the exercise toward the channels that are easiest to track rather than the ones that work.

The better approach models each source at the grain it’s available and lets the layers inform each other. Aggregate data constrains the totals; user-level data, where it exists, identifies the mechanisms — sequence, frequency, creative interaction — that aggregate data can only imply. Bayesian methods handle this well: coarse data acts as a prior on the fine, and a channel with thin data ends up with wide confidence intervals rather than a confidently wrong point estimate. The consequence is that a brand doesn’t have to wait for a perfect data foundation. Better coverage improves the model; it isn’t a precondition for having one.

Where Causal Modeling Fits In

It’s tempting to apply causal modeling wherever possible, but there’s a time and a place for it. Causal inference depends on separating signal from noise, which takes a significant volume of observations. Small advertisers generally don’t have it; high-volume national advertisers do — and the more they want to slice the answer by SKU, region and audience, the more volume each slice needs. Coverage matters as much as scale: a brand spending heavily inside a single walled garden can get channel-specific insight, but can’t see how channels affect each other without a model that spans them.

None of this replaces the layers above or below it. The MMM still sets the strategic envelope, and bidding and daily optimization still happen in the platforms. Causal modeling is the layer in between that actually helps you understand what to adjust and what the downstream impacts of those adjustments might be.

The Penn District