Skip to content
← Back to Blog
·9 min read·Hass Dhia

Outcomes-Based Measurement Needs a Stress Test. Hallmark's iSpot Deal and the S&P 500's Worst Days Show Why.

outcomes-based measurementmarketing measurementbehavioral economicsdecision intelligencebrand strategy

Picture a commercial for Hallmark's upcoming movie Holiday Ever After: A Disney World Wish Come True. A viewer sees it, picks up a phone, and starts booking a Disney vacation. That sequence presumably happens all the time. What Hallmark is now offering is a way for advertisers to see it, at scale, across linear, digital, and FAST channels.

On October 1, Hallmark Media announced a partnership with measurement company iSpot to do exactly that. Adweek's report puts the promise in a single line: when commercials around the Disney movie send people online to book a trip, "advertisers can finally know."

Outcomes-based measurement is the right direction. It also inherits a problem from a field that has had longer to find it. A Kiplinger piece published the same day shows how one true statistic can argue for opposite conclusions depending on which half of the data you show. Outcomes numbers will do the same thing unless buyers learn to ask for the other half.

What Hallmark Is Actually Selling

The announcement reads like a technology story. It is more interesting as a change in who carries the burden of proof.

According to Adweek, the deal lets Hallmark link cross-platform ad exposure to real-world actions: website visits, box-office sales from movie-trailer exposure, and in-store purchases. iSpot's executive vice president of media partnerships, Stuart Schwartzapfel, describes the pitch as "true unification of Hallmark's linear, digital, and FAST properties, all measured through one lens." The measurement also covers deduplicated reach and incrementality, the two features meant to stop an outcome number from counting the same viewer twice or crediting an ad for what would have happened anyway.

It arrives with a credential attached. The Joint Industry Committee recertified iSpot as a national currency for the 2026-2028 term, and an iSpot executive told Adweek that gives clients additional confidence. Hallmark's Casey Gould, senior vice president of ad sales and advanced advertising, was candid about where this stands: "We're in the first inning."

Who Carries the Burden of Proof

Notice what moved. Ratings measure delivery. Outcomes measure effect. Shifting a sales conversation from the first to the second changes what a seller can be wrong about, and Hallmark is offering to be judged on the harder currency. The Adweek report does not say who pays iSpot, which matters for how much weight "independent" deserves, in the same way it matters in any audit.

There is a second limit. Outcome data can show that something worked. Explaining why, and predicting when it will stop working, is a separate job. We covered that gap in our look at why creator content converts, where marketers could see the conversion and still could not explain it. Seeing an outcome and understanding it are different assets, and the second is the one a budget defense needs.

The Statistic That Argues Both Sides

The Kiplinger article was written by Jordan Rizzuto, managing partner and chief investment officer at GammaRoad Capital Partners. It tests a refrain every investor has heard: "Missing only the 30 best days of market returns can meaningfully lower your portfolio value."

The test uses the S&P 500 Total Return Index from January 4, 1988 through July 31, 2026. A buy-and-hold investor earned 11.46% annualized, turning $1 into $65.48. Miss the 30 best days and the return falls to 6.32%, leaving $10.60, roughly 84% less wealth. That is the version most people know.

Now delete the 30 worst days instead. The same dollar compounds at 17.41% a year and ends at $487.56, more than seven times the buy-and-hold result. Same index, same period, same arithmetic, opposite moral.

What the Tails Actually Show

Two details make the picture sharper. The article notes that these extremes are rare, "just 30 out of 9,716 market days," about 0.3% of the sample. And they do not scatter randomly: since 1988, a best or worst day has fallen within 21 trading days of another one more than 70% of the time, and nearly all of them arrived during bear markets, when the S&P 500 had already dropped 29.17% from its prior peak on average.

Here is the part the article does not state. The dollar gap looks lopsided, but most of it is the arithmetic of multiplying up versus dividing down from the same starting point. Deleting the worst 30 days multiplies the buy-and-hold result by about 7.4 ($487.56 against $65.48). Deleting the best 30 divides it by about 6.2 ($65.48 against $10.60). Measured in annualized return, the two edits move the figure 5.95 points up and 5.14 points down. The author says the worst days carry "far greater influence" on annualized returns, yet the article's own figures put the difference between the two tails at 0.81 points a year. The tails are roughly symmetric. When the two halves of a statistic are that close, the choice of which half to show is doing much of the persuading.

The article's stronger argument is a different one. A 10% loss needs an 11.11% gain to get back to even, a 25% loss needs 33.33%, a 40% loss needs 66.67%, and a 50% loss needs 100%. For an investor with a real deadline, such as retirement in five years or a tuition bill, the time it takes to climb out of a drawdown is the cost. That argument is about path and deadlines. It is sound, and it has little to do with deleting thirty days from a historical series.

The Author's Interest and the Concession

Read the byline too. The author is chief investment officer at GammaRoad, whose bio says its systematic strategies seek to improve portfolios by "avoiding the worst market drawdowns." That is worth knowing when weighing the 17.41% case, which is the argument for drawdown avoidance. It is also a ghost, because nobody can dodge exactly the 30 worst days, and the article concedes as much: the question worth asking "is not whether you can time the best or worst days."

That concession is the useful part. If you cannot time the days, then choosing which tail to delete is a rhetorical decision, not a strategy. Delete the best days and you argue for patience. Delete the worst and you argue for risk management. Both are arithmetic. Neither is advice, and nothing here is investment advice either. The sentence worth keeping is the author's own: "The question to ask is whether your portfolio is built to withstand the market environments where the worst days tend to occur."

Marketing Has Its Own Thirty Days

Carry that structure over to outcomes-based measurement. Imagine an outcomes report that foregrounds the best days of a campaign, the exposures followed by an action. Delete those and little is left, the marketing equivalent of the 84% loss. The number would be real. It would also be half the picture.

Nothing in the Adweek report shows that advertising outcomes cluster the way market returns do. That is the first thing a tail test should establish. But two facts from the report make the question worth asking. Hallmark's strongest ratings are seasonal: it "regularly dominates Q4 TV ratings," in Adweek's words, and Countdown to Christmas starts October 16. If outcomes concentrate with the season, a year's measured value rides on a few months, and the instinct to trim spend in a soft stretch becomes the marketing version of selling into a drawdown.

Where the Analogy Breaks

The analogy has limits worth stating. Market returns compound along a single path, while campaign outcomes tend to add across placements, so the clustering and compounding of investing need not carry over. What carries over is the disclosure habit, not the arithmetic. Hallmark's deal already includes deduplication and incrementality, which address whether an ad changed behavior on average. The tail test asks a different question: how that effect is distributed across days, placements, and conditions, and what happens to the decision when it fails.

Three questions turn a measurement claim into a tested one.

The best-days question

What share of the measured outcomes came from the top 1% of days, placements, or audiences? If most of a campaign's lift lives in a few moments, the average says little about what a marginal dollar does. A concentrated result is not a bad result. It is a result that should change how the budget is paced.

The worst-days question

What share of spend landed where nothing measurable happened, and does the report show it? A dashboard that leads with lift and buries dead spend shows the middle of the distribution and hides the tail that decides budgets. In our analysis of Gap's decade of decline, the mechanism that defunded brand investment was motivated reasoning rather than financial conservatism. A report with no tail disclosure hands that reasoning an easy excuse: one bad measured quarter, and no context for how bad quarters are distributed.

The regime question

Do the results hold in the environment where budgets get cut? Kiplinger's data say extreme days arrive in bear markets. Campaign measurement should be stress-tested the same way: run the outcome model on the weak quarter, the unfamiliar market, the audience that does not already love the brand. We argued earlier that fandom metrics are survivorship bias in disguise, a measurement trap that flatters a brand with the customers it already has. Outcomes measured only on loyal households inherit the same flaw.

Stress Tests Show Up Wherever the Cushion Disappears

Branding Strategy Insider makes the same argument from a different field. Its guest author, writing about international hiring, says companies do not necessarily need a new employer brand for every market: "You need to stress-test the one you already have." At home, a strong culture can cover for a clumsy offer letter or a slow payroll run, because people experience the organization around them. At a distance the cushion disappears. The author imagines a candidate in Lisbon or Nairobi who reads the mission statement, accepts an offer, and then waits eleven days for a contract. The brand has already changed in that person's mind.

The line that matters is the diagnosis: "Expanding into new markets does not necessarily create employer brand problems. It exposes them." Distance, the author adds, "removes the informal goodwill that used to hide them." Put the three fields side by side. A bear market strips the cushion from a portfolio. A border strips it from an employer brand. Outcome measurement, done properly, strips it from a media plan, by asking what an ad did when no brand halo, seasonal tailwind, or loyal audience was doing the work.

The same article contains the best argument for what Hallmark just bought. It calls a specific promise like giving small teams real ownership from week one "valuable precisely because it is testable." An advertiser's promise becomes testable when outcomes are measured. The risk is that a testable claim gets tested only on its best days.

The Macro Backdrop

The macro backdrop adds pressure. McKinsey's executive summary of its August 2026 Global Economics Intelligence briefing says US debt has hit a record high and that inflation fueled by energy supply uncertainty "limits room for maneuver on interest rates." This section relies on that published summary, not the full report. The reading is modest: when policy has little slack, companies lose the option of waiting for conditions to repair a fragile plan, and finance teams lose patience for spend that cannot show its work. The sensible response is robustness designed in advance, which is exactly what the Kiplinger author and the employer brand author each prescribe in their own field.

What to Ask Before You Believe the Number

Outcomes-based measurement will spread, and it should. A buyer holding one of these reports can run the tail test in an afternoon.

Ask for the concentration: where did the outcomes come from, and how few days or placements carry the result? Ask for the dead spend: what share of the budget produced nothing measurable, shown next to the lift and not in an appendix? Ask for the weak-period result: what happened in the quarter, market, or audience where the brand had no cushion?

The timing favors buyers. Gould told Adweek that Hallmark's offerings to advertisers on iSpot are set to grow, which suggests the reporting templates are still being written. Requests for concentration and dead-spend views are cheap to make before the first template hardens into the default, and expensive to retrofit once every campaign report in the category looks the same.

What Measurement Is For

The conclusion is less comfortable than the sales pitch. The point of measuring outcomes is not to prove that advertising works. It is to learn the conditions under which it stops. A number that cannot name its own failure conditions is a best-days statistic, and the investing example shows how a best-days statistic can be true and still mislead.

Every source here has something to sell, as sources do, and that includes this one. The defense is the same one a reader should demand from any of them: a number should arrive with its method. That is also the premise behind STI's open-source research, which lives on our research page. Ask for the tails before you sign the budget.

Want more insights like this?

Follow along for weekly analysis on brand strategy, market dynamics, and the patterns that separate signal from noise.

Browse All Articles →

Or explore partnership opportunities with STI.

Related Articles