The Burnt-Money Signal
A Super Bowl ad costs ~$7M for 30 seconds plus production. A surprising number of them convey no verifiable information — a celebrity, a feeling, a logo. The puzzle: why does burning money on a content-free message work? The answer is that the burning is the message. Cost is the medium.
A firm knows its own quality; the buyer does not. Let quality be a type:
- $\theta \in {H, L}$ — High or Low. The firm knows $\theta$; the buyer sees only the prior $\Pr(H) = p$.
- The product is an experience good: quality is revealed only after purchase (and often after repeat purchase). The buyer can’t inspect it at the point of decision.
This is the precondition. Signaling theory has bite precisely where the buyer “can’t otherwise verify.” For a search good (quality visible pre-purchase), you’d just show the product and the whole apparatus collapses.
Why cheap talk fails first
Before building the expensive solution, see why the cheap one doesn’t exist.
Suppose the firm could just say “we’re high quality.” A message with zero cost imposes the same payoff regardless of type. So whatever a High firm says, a Low firm can say identically at the same (zero) cost. The buyer knows this, so the message carries no information — it can’t move the posterior off the prior.
Formally, cheap talk (Crawford–Sobel) only transmits information when the sender’s and receiver’s preferences are partially aligned over outcomes. Here they’re directly opposed on the thing being communicated: both types want the buyer to believe “High.” With fully opposed interests on the state, the unique outcome is the babbling equilibrium — all messages are noise, the buyer ignores them. The ad’s words (“Tastes great,” “Believe”) are exactly this babble. That’s the clue: the content was never doing the work.
The costly action
Replace “saying” with “doing something expensive.” Let the firm choose advertising spend $a \ge 0$. The key is that the cost of the action differs by type. Let the firm’s payoff be:
$$
U(\theta, a, b) ;=; R(\theta)\cdot b ;-; c(\theta), a
$$
where $b \in {0,1}$ is the buyer’s belief/purchase response (buy if convinced High), $R(\theta)$ is the revenue captured if the buyer buys, and $c(\theta)$ is the firm’s marginal cost of advertising effort.
The crucial structural assumption — the single-crossing / Spence–Mirrlees condition:
$$
c(H) ;<; c(L).
$$
The High type pays less per unit of signal. Why would that be true for an ad that costs the same dollars to air?
Because the ad is an investment recouped only by quality. The mechanism isn’t the airtime price — it’s that the spend pays for itself through repeat purchase and word of mouth, which only materialize if the product is actually good. A High firm spends $7M to acquire a customer who comes back; the lifetime value covers the burn. A Low firm spends the same $7M to acquire a customer who tries the product once, is disappointed, never returns, and tells others. Same nominal outlay, but the Low firm has no back end to amortize it against. Net cost per dollar of ad is higher for the Low type. The Super Bowl ad is a bond the firm posts against its own future, redeemable only if buyers like what they get. (Nelson 1974; Milgrom–Roberts 1986 formalize this as “dissipative” advertising that signals through expenditure, not content.)
This is why the ad can — even should — say nothing about the product. The information is in the magnitude and visibility of the spend, not the claims. A verifiable claim would invite a verifiable rebuttal; the burn needs no adjudication. Conspicuousness is itself part of the technology: the Super Bowl is the most expensive, most-watched slot precisely because everyone knows it’s the most expensive — the cost is common knowledge, which is what a signal requires.
Separating equilibrium
A separating equilibrium: the two types choose different actions, so the buyer perfectly infers quality.
Let $H$ spend $a^* > 0$ and $L$ spend $0$. The buyer’s belief: “spend $\ge a^* \Rightarrow$ High, buy.” For this to hold, two incentive-compatibility constraints must bind:
(IC-L) The Low type must not want to mimic. Mimicking buys the “High” response $R$ but at Low cost-per-unit:
$$
R(L) ;-; c(L),a^* ;\le; R(L)\cdot \Pr(\text{buy} \mid 0).
$$
If non-signaling yields zero sales, this is $R - c(L),a^* \le 0$, i.e.
$$
a^* ;\ge; \frac{R}{c(L)}.
$$
(IC-H) The High type must still prefer to signal:
$$
R ;-; c(H),a^* ;\ge; 0 \quad\Longrightarrow\quad a^* ;\le; \frac{R}{c(H)}.
$$
A separating $a^$ exists in the interval
$$
\boxed{;\frac{R}{c(L)} ;\le; a^ ;\le; \frac{R}{c(H)};}
$$
which is non-empty exactly because $c(H) < c(L)$. Single-crossing is what carves out room for a spend that High will pay and Low won’t. The signal works not despite being expensive but through being expensive enough that the cheater can’t afford to fake it.
Note the spend is purely dissipative — it produces no direct utility, conveys no facts. Burned money. Its entire value is informational: it makes belief self-confirming.
Pooling equilibrium
A pooling equilibrium: both types choose the same action, so the buyer learns nothing and stays at the prior $p$. Two flavors:
- Pool on silence ($a = 0$ for both). Sustainable when the prior $p$ is high enough that buyers purchase anyway — no one needs to signal. (A category where average quality is trusted.)
- Pool on spend — both advertise. Sustainable only off the equilibrium path with pessimistic beliefs (“anyone who doesn’t spend is Low”), which is usually fragile.
Equilibrium selection sorts this out. Pooling equilibria typically fail the Intuitive Criterion (Cho–Kreps): ask “which type could credibly deviate?” A High type can deviate to a spend that a Low type would never rationally choose even if it relabeled them High — i.e., an $a$ in the separating band. Since only $H$ benefits from that deviation, the buyer should read the deviation as “High.” That unravels the pool. The refinement pushes us toward the least-cost separating equilibrium: $a^* = R/c(L)$, the cheapest spend that still keeps Low out. Even efficient signaling burns money; it just burns the minimum that deters mimicry.
Why the waste is the point
The efficiency story is genuinely double-edged:
- Socially, the spend is partly deadweight. The dollars buy no product improvement. Spence’s original insight (education as signal) carries over: signaling can be individually rational and collectively wasteful.
- But it’s not pure waste when the asymmetry is real and severe. The signal lets High firms enter and survive a market that would otherwise suffer a lemons collapse (Akerlof) — if buyers can’t tell good from bad, they pay average price, High firms exit, quality spirals down. Costly signaling is the institution that prevents the unraveling. The burn is the price of credibility in a market with no cheaper verification.
The model in one breath
Cheap talk fails because both types want to claim “High” at zero cost, so words are babble. A costly action separates them only if the High type’s cost per unit of signal is lower — which holds for advertising because the spend is recouped through repeat purchase that only good products earn. The equilibrium spend sits in the band $[R/c(L),, R/c(H)]$: high enough that a Low firm would rather not bother, low enough that a High firm still profits. The ad therefore says nothing about the product on purpose — its content is irrelevant; its cost, made common knowledge, is the entire message. A firm that can afford to set $7M on fire in front of 100 million people is announcing it expects you to come back.
A few load-bearing caveats worth flagging, since the model is clean and reality isn’t:
- The $c(H) < c(L)$ ordering is the whole game. If a Low firm can also profitably advertise (e.g., a one-shot purchase with no repeat-business back end, or a scam that exits before reputation catches up), single-crossing breaks and the signal stops separating. Empirically this is why signaling theory fits durable, repeat-purchase, brand goods better than one-off purchases.
- Real Super Bowl ads are rarely pure burn — they also build awareness and memory (a direct demand effect), which muddies the clean “dissipative signal” reading. The pure model isolates the signaling channel; attribution between “I now know this exists” and “I now infer it’s good because they could afford this” is an empirical question, not settled by the theory.
If you want, I can push this further in any of three directions: (1) work the algebra for a continuum of types with a smooth $c(\theta)$ and derive the separating spend schedule $a(\theta)$; (2) add repeated-game microfoundations so $c(H)<c(L)$ falls out of a reputation model rather than being assumed; or (3) contrast with money-burning in cheap-talk games (Austen-Smith–Banks) where burning can transmit info even without a type-dependent cost.
Here is a formal game-theoretic analysis of the Super Bowl advertising expenditure through the lens of Costly Signaling.
In the Standard Game of Product Quality, there is Hidden Information.
- The Firm (Sender): Knows the true quality $q$ of the product ($q \in {q_H, q_L}$).
- The Consumer/Market (Receiver): Only knows the prior probability that a firm is High Quality.
- The Market Failure: Without a signal, consumers cannot distinguish $q_H$ from $q_L$. They only pay the Average Quality Price ($P_{pool} = q_{expected}$). This leads to “Cream Skimming” or the “Market for Lemons” problem, where Low Quality firms undercut High Quality firms until the market collapses to the low-quality reservation price.
To solve this, High Quality firms must send a credible signal. If communication costs are zero, Low Quality firms will simply lie (Cheap Talk). To prevent pooling equilibria where liars deceive, the signal must be Costly.
2. Game Setup: Parameters & Payoff Functions
Players:
- Firm: Type $\theta \in {H, L}$ (High Quality, Low Quality).
- Consumer: A representative indifferent buyer.
Private Information:
- Quality: $q$. $q_H > q_L$.
- Prior Belief: Probability of High Quality Type $Pr(\theta=H) = \pi$.
Strategies:
- Signal ($s$): Public Advertising Expenditure. $s \in \mathbb{R}, s \ge 0$.
- Consumer Action ($b$): Buying Decision / Willingness to Pay ($p$).
Payoff Structure:
The Firm’s utility (Profit $\Pi$) is the revenue from consumers minus production cost and advertising cost.
$$
\Pi(\theta, s, p) = p(q) - C_{prod}(\theta) - C_{adv}(s, \theta)
$$
The Crucial Cost Function Assumption:
For signaling to work, the cost of sending the High Signal ($s_H$) must be manageable for the High Type but unsustainable for the Low Type.
- Production Cost: Assumption: High Quality firms have higher efficiency or durability, allowing higher profit margins ($q$ is private, but correlates to $q$‘s viability).
- Advertising Cost: $C_{adv}(s, \theta)$. Crucially, we assume the cost is nonlinear and type-dependent.
$$
C_{adv}(s, \theta) = s \cdot (1 + k\theta)
$$
Explanation: For the Low Quality firm ($L$), the “multiplier” is high. Even if $s$ is the same for everyone, the ability to sustain the revenue stream generated by that burn is lower for $L$. Alternatively, we can define the cost as a Sunk Risk Cost:
- High Quality ($H$): The ad burn funds reputation that compounds over time ($V_H$ is high).
- Low Quality ($L$): The ad burn funds reputation that fails instantly ($V_L$ is low or negative).
Simplified standard Spence formulation for this context:
$$
\text{Cost}_H(s) < \text{Cost}_L(s) \quad \text{for the same signal intensity } s
$$
(In the marketing context, this means: For a given brand investment, the High Quality firm gets a higher Return on Investment (ROI) and lower long-term churn risk than the Low Quality firm.)
3. Why “Cheap Talk” Fails
Cheap Talk is communication with zero cost ($s = 0$).
Equation: $\Pi = p(q) - 0$.
In a Cheap Talk game:
- Incentive to Lie: Both types $H$ and $L$ want to be perceived as High Quality to charge price $p_H$.
- Consumer Risk: Believing the Low Type claims $q_H$ is rational until the product is consumed.
- Reputation Cost: In a one-shot game, Low Type firms lie with probability 1. Consumers learn this quickly and revert to $p_L$.
- Second-Stage Credibility: Because there is no cost to the lie, $L$ says “We are High Quality” regardless of quality. The signal provides zero information.
$$
\mu(q| “We are High”) = \text{Prior Belief}
$$
- Result: Both High and Low types end up paying the floor price $p_L$. The High Quality firm loses money.
Why Super Bowl Ads Work (vs. Cheap Talk):
The ad is not just a claim; it is a financial sacrifice.
- Spending $5M is irreversible.
- If a Low Quality firm sends the signal $s_{high} = $5M$:
- Consumers pay $p_H$ (boosted by the ad).
- Product Quality $q_L$ is revealed upon consumption (high failure rate).
- Reputation Collapse: The stock price crashes, and future churn costs exceed the ad revenue.
- The Cost: The Low Quality firm goes bankrupt immediately.
- For the High Quality firm: The churn is low, brand equity is high. The sunk cost is recouped via high-margin future sales.
4. Equilibrium Analysis
A. Why the Ad Says “Nothing” About the Product
Psychological Signaling: By not specifying specs (MPH, MPG, Gigabytes), the advertiser forces the consumer to rely on Brand Trust.
- Verifiable qualities (e.g., “Battery lasts 20 hours”) can be faked by inventing specs.
- Emotional qualities (e.g., “Family”, “Freedom”) are unadministered qualities. They signal, “We have the cash to buy time with you.”
- The Signal is Behavior: It says, “We are willing to write a check you cannot see into the balance of our future cash flows.”
B. Separating vs. Pooling Equilibrium
1. Pooling Equilibrium (All say High):
- Strategy: $H$ sends $s=5M$, $L$ sends $s=5M$.
- Belief: Consumers believe average quality: $q = \pi q_H + (1-\pi) q_L$.
- Price: $P_{pool} = q$.
- Profit for High: $\Pi_H = q p_{pool} - C_{prod}(H) - C_{adv}$.
- Profit for Low: $\Pi_L = q p_{pool} - C_{prod}(L) - C_{adv}$.
- Failure: Because $C_{adv}$ (the risk of future failures) is higher for Low, they eventually cannot sustain the signal. If $L$ enters Pooling, they create a bad reputation. Furthermore, High Quality firms will pay to exit the ad spend to get the high price, provided they can signal without the risk.
- Why it fails: Low quality firms prefer to cheat unless they face the penalty of failing the reputation balance.
2. Separating Equilibrium (Signal is Credible):
- Strategy:
- $\sigma_H(s=5M) = 1$ (High Ad Spend).
- $\sigma_L(s=0) = 1$ (No Ad Spend / Low Entry).
- Belief:
- If $s = 5M$: Belief $= q_H$ (Consumers pay $P_H$).
- If $s = 0$: Belief $= q_0$ (Consumers pay $P_0$ or reject).
- Payoff for High:
$$
\Pi_H(Separated) = P_H (q_H) - C_{prod}(H) - C_{adv}(5M) > P_{pool} (q_H) - C_{prod}(H)
$$
(High Quality benefits from the price premium more than it pays for the ad burn).
- Payoff for Low:
$$
\Pi_L(Separated) < P_{pool} (q_L) - C_{prod}(L)
$$
(Low Quality cannot profit from $P_H$ because $s=5M$ is too costly/risky $> P_H$ margin).
3. The Separating Condition:
For the equilibrium to hold, the Low Quality firm must be “profit-drained” by the signal:
$$
C_{adv}(s, L) \ge \text{Excess Profit of High Type from } q_H \rightarrow \text{Minus } q_L \text{ Disutility}
$$
Essentially:
$$
\text{Cost of Burning $5M} > \text{Marginal Gain from being perceived as High Quality (for L-type)}
$$
And:
$$
\text{Cost of Burning $5M} < \text{Marginal Gain from being perceived as High Quality (for H-type)}
$$
This is only possible if $L$-types are effectively raiding their future survival to buy a present-day reputation drop. High Quality firms are not raiding survival; they are optimizing value.
5. Why the “High Quality” Model demands millions in Abstract Ads
Under this game theoretic framework, Super Bowl ads answer two specific questions:
- Liquidity Test: Can you raise $5M without borrowing? Companies that borrow $5M for ads are hiding financial stress. High Quality companies have the cash on hand.
- Sunk Cost Fallacy Prevention: The ad is designed to be abstract (“We Love Our Customers”) so it cannot be refuted by ticking a technical box.
- If the ad said “Our Engine is 10% better,” you can verify it.
- If the ad said “This car belongs to those who see what others miss,” you cannot verify it easily unless you attend the product launch.
- This creates a High Incentive Signal. High Quality firms have the confidence to claim they offer value beyond the product. That value is the ability to predict markets, not the product specs. Low Quality firms cannot claim “value beyond product” because their product is the only thing they have, and they lack the product longevity.
Conclusion
The Super Bowl ad is a Costly Signal in a Bayesian Game.
- Cheap Talk ($S=0$): Fails because the “cost” is low to fake.
- Separating Equilibrium ($S>0$): High Quality firms pay the cost because their Private Information (High $q$) allows them to reap revenue that covers the sunk cost.
- Low Quality Firms: Cannot mimic this without incurring a loss of market share/reputation/bankruptcy.
The millions spent are not for the buyer’s eyes; they are a “Proof of Capital” requirement to enter the cooperative game with consumers who cannot see the quality. It forces the “Ideal Type” to reveal itself by the magnitude of the burn, separating it from the “Lemons.”
Where the cost asymmetry actually lives (read before the model)
The common setup is wrong in an instructive way: the Super Bowl ad does not signal because the ad is more expensive for a low-quality firm to run. The network charges every buyer the same posted price for the slot, independent of quality. The ad cost is identical across types, so the signal cannot work on the cost side the way Spence’s education does (where the high type genuinely finds the signal cheaper to acquire). Single-crossing must therefore be located on the benefit side.
The benefit-side mechanism (Nelson 1974, Advertising as Information, JPE 82(4):729–54; formalized by Milgrom & Roberts 1986, JPE 94(4)): the same ad dollar earns a higher return for the high-quality firm because its product generates repeat purchases, while the low-quality firm gets one disappointed trial and no comeback. Single-crossing lives in payoffs/benefits, not costs.
“Says nothing” is load-bearing, not incidental. A content-free ad is pure dissipative money-burning, which strips away every channel of value except the inference “only a firm expecting repeat business would set this money on fire.” Adding hard product claims is optional; the burn is the signal. This is what makes it the clean case.
Three Phase A assumptions worth flagging so you can correct them: (1) a formal signaling game is wanted, not just intuition; (2) benefit-side single-crossing is the right primary mechanism for a content-free ad — switch to cost-side if your intended meaning is “prestige production values low-quality firms can’t match”; (3) the comparison target is a babbling cheap-talk equilibrium.
Players and payoffs
Nature (chance node, Move 0) — draws firm quality θ ∈ {H, L}, prior Pr(H) = p. How inferred: structural-position (the modelling device that makes the information asymmetry concrete).
Firm (Sender, informed) — Actual value terms (not claimed-to-want): maximizes lifetime/multi-period profit net of ad cost. How inferred: revealed-by-behaviour. Where actual diverges from claimed: the firm claims to want “to inform consumers about our great product,” but the firm does not care whether buyers learn the truth — it cares whether they buy. A low-quality firm would happily deceive if deception paid. This divergence between claimed and actual payoff is exactly what the low type would exploit if the signal were cheap, and exactly why cheap talk fails.
Consumers (Receiver, uninformed) — Actual value terms: maximize net consumption surplus — positive from an H purchase, negative from an L purchase, zero from not buying. How inferred: revealed-by-behaviour. Where actual diverges from claimed: they claim to want “to find good products,” but they do not reward honesty; they reward being right.
Firm-side payoff structure (illustrative numbers; only the orderings are load-bearing). Ad cost C identical for both types. Gross value realized only if consumers then buy: high type induces trial → satisfaction → repeat business → high return; low type induces one purchase → no repeat → low return; not advertising → 0. One coherent instantiation: C = 7; high-type return = 20 (or, in an explicit two-period frame, 2π = 10 where per-period profit π = 5); low-type return = 5 (= π). The two numeric instantiations across this analysis are illustrative variants of the same ordering, not a substantive disagreement.
Resulting firm net payoffs (two-period instantiation): H: Ad = 2π − C = +3, No Ad = 0. L: Ad = π − C = −2, No Ad = 0. The row gap is larger for H than for L only because H collects the second-period repeat stream — this is the single-crossing condition, met through benefit asymmetry (only H earns the repeat return on the identical-cost burn), not Spence’s cost asymmetry.
Consumer-side payoffs (instantiated so the receiver’s node is derived, not asserted). Surplus from buying H = s_H > 0; loss from buying L = s_L < 0 (you paid for quality you didn’t get); not buying = 0. One instantiation: s_H = +10, s_L = −8; an equivalent variant: s = 4 surplus / d = 6 loss. A consumer with belief μ = Pr(H | action) buys iff μ·s_H + (1−μ)·s_L ≥ 0 ⟺ μ ≥ p* ≡ −s_L/(s_H − s_L) (= 8/18 ≈ 0.44 in the first instantiation).
The load-bearing orderings (the entire result): V_L < C < V_H (i.e. π < C ≤ 2π) and s_L < 0 < s_H. The illustrative magnitudes do not matter.
Missing-player flag: the slot seller (NFL / network), the regulator (FTC), and third-party certifiers are reactive parties whose responses would shift the equilibrium. Their inclusion is recommended where their behaviour is observable; their absence is named here so the equilibrium below is read as bounded-to-the-current-inventory. They are detailed at the end of this section and revisited as alternative structures.
The fuller inventory of reactive players, each modeled here as a parametric channel rather than a fully endogenized player:
- Regulator (FTC / false-advertising law) — the most important omission. Legal penalties for false claims attach a cost to lying, converting cheap talk into costly talk; with a large enough penalty, “our product is high quality” becomes self-enforcing and the firm no longer needs to burn the millions. Regulation and dissipative signaling are substitutes, not complements. (Bounded: enforcement probability/intensity not modeled — would require a chance node.)
- The slot seller (NFL / network) — C is endogenous, not exogenous. A monopolist on the scarce credible-burn medium can capture the signaling rent by pricing the slot up toward 2π − ε (just below where even H walks away), squeezing H’s net surplus toward zero and migrating the rent from advertiser to medium. The observed price climb ($7M → $8M → $10M) is consistent with a monopolist raising price toward H’s reservation value as audiences fragment and scarcity rises.
- Rival / competing high-quality firms. If many H firms must advertise to separate, the result is a dissipative arms race (Red-Queen cost of pooling-avoidance) — each H spends C, profits compete away, total advertising becomes a social cost the two-player frame understates. A low-quality competitor can also free-ride on category demand the ad stimulates.
- Third-party certifiers / review platforms. A cheap verification channel strips the dissipative premium from the ad spend (leaving the awareness floor) — a genuine disruption risk to the burn-as-signal strategy.
Explicit boundedness flag. The slot seller, regulator, and certifier are modeled here as parametric channels, not fully endogenized players with their own equilibrium strategies (only the slot seller gets a payoff sketch). Endogenizing the seller’s pricing and the certifier’s entry would materially change welfare and the size of the optimal burn — flagged as out of the current frame.
Game classification
Timing: sequential / extensive-form. Nature → Firm chooses {Ad, No Ad} (more generally a burn level) → Consumers observe the action only and choose {Buy, Don’t}. A dynamic signaling game, not simultaneous.
Information: incomplete and imperfect. Incomplete: nature’s type draw is private to the firm (Harsanyi type, prior p). Imperfect: consumers act at an information set that does not reveal θ — they see the burn, not the quality.
Duration: one-shot signaling stage embedded in a repeated/multi-period product market. The signal is sent once; purchasing repeats. The mechanism (high-type return > low-type return) exists only because the firm–consumer relationship repeats — repeat purchases are what the high type monetizes. The repeat structure is intrinsic, not decorative (tested under alternative structures).
Sum: mixed-motive, positive-sum on the equilibrium path. When H signals and consumers buy, both gain (separation creates value by matching consumers to quality). On the low type, interests conflict: the firm wants the sale, the consumer doesn’t. Not zero-sum.
Equilibrium analysis
Equilibrium method: Perfect Bayesian Equilibrium, refined by the Cho–Kreps (1987, QJE) Intuitive Criterion. PBE requires: (a) each type’s action is sequentially rational given consumer strategy; (b) consumer strategy is sequentially rational given beliefs; (c) beliefs are Bayes-consistent with strategies on the equilibrium path.
Signal space stated explicitly. The firm chooses a burn level e ∈ [0, ∞) at cost e (identical across types — money is money); gross value is realized only if consumers then buy. The binary {Ad = C, No-ad = 0} set used in the separating derivation is a coarse discretization of this continuum, adequate to implement separation when C lands in the band; the continuum is what the Cho–Kreps refinement argument requires, so it is declared up front rather than smuggled in.
Pessimistic-prior / lemons regime is the operative case. With the instantiated numbers, expected surplus from buying blind is negative (e.g. p·s_H + (1−p)·s_L = 1.6 − 3.6 = −2 < 0, or equivalently p < p* ≈ 0.44), so consumers will not buy without a signal. This is the economically meaningful regime in which the Super Bowl burn earns its keep. (When p ≥ p* / optimistic prior, consumers buy at the prior and no signal is needed.)
Derivation — separating equilibrium (H burns, L doesn’t). On-path beliefs μ(Ad) = 1, μ(No Ad) = 0. Consumer best response, derived from the threshold: μ(Ad) = 1 > p* ⇒ Buy after Ad; μ(No Ad) = 0 < p* ⇒ Don’t buy after No Ad. Type checks: High type — Ad → 2π − C = +3 > 0 ⇒ advertise ✓. Low type — Ad → π − C = −2 < 0 ⇒ stay silent ✓. The low type’s deviation succeeds at deception (consumers do buy) and still loses money, because it cannot generate the repeat business needed to recoup the burn off one-shot sales. That is the whole engine. Beliefs are Bayes-pinned on path; no off-path freedom needed.
Existence condition (stated once, consistently): V_L < C ≤ V_H (i.e. π < C ≤ 2π) — low type strictly declines (C > V_L ⇒ imitation loses money), high type weakly advertises (C ≤ V_H ⇒ signaling still profitable). The cost must sit in the separation band: high enough that the low type can’t recoup it, low enough that the high type still profits. Knife edges (C = V_L low type indifferent; C = V_H high type indifferent) are measure-zero and don’t change the prediction. The least-cost separating outcome (Cho–Kreps selection) puts H’s burn just above V_L (e* → V_L⁺).
Sharpened separation condition (take-the-money-and-run). “π” must be read as the maximum attainable one-shot fooling revenue for L, π′, not a fixed number: a well-capitalized low-quality firm that converts the ad’s awareness reach into a one-time sales spike defeats separation if π′ > C even with zero repeat stream. The load-bearing condition is therefore π′ (max one-shot fooling revenue for L) < C ≤ 2π (two-period honest return for H). This is part of why the credible-burn medium must be so expensive (the slot must be priced above the largest one-shot fooling haul), and it locates the live exploit: deep-pocketed, scalable low-quality entrants.
Both pooling configurations enumerated and defeated.
- Pooling-on-advertise: in the lemons regime the pooled belief μ = p < p* fails the buy threshold, so consumers wouldn’t buy on the pooled path; both types would have burned for nothing and strictly prefer e = 0. Equivalently, L’s pooled payoff π − C = −2 < 0 ⇒ L strictly deviates to No Ad. Pool-on-Ad cannot survive — the ad is too expensive for L even when pooled. (It can exist only when p ≥ p*, the uninteresting no-signal-needed regime, dominated there by everyone burning nothing.)
- Pooling-on-no-ad: neither type burns; μ(No Ad) = p. Sustainable as a PBE only under pessimistic off-path beliefs (treat any surprise burn as low-type) or when the prior is optimistic enough that consumers buy anyway (p·s_H ≥ −(1−p)·s_L; e.g. p = 0.7 → +1 > 0, an established brand / standing reputation). In the pessimistic case the Intuitive Criterion kills it: for a deviation burn e′ ∈ (V_L, V_H), the low type’s best conceivable payoff V_L − e′ < 0 is worse than its pooling 0 (equilibrium-dominated for L), while the high type’s V_H − e′ > 0 is not dominated — so consumers must assign μ(H | e′) = 1, buy, and the high type profitably deviates. Pooling-on-no-ad unravels, leaving the least-cost separating equilibrium as the unique survivor. (Milgrom–Roberts 1986 show the same Intuitive-Criterion logic “rules out all pooling equilibria” in the advertising setting; web- and Kellogg-corroborated.)
Dissipative-signaling result. μ updates on the action {Ad, No Ad} and its cost C, not on words. The decoded sentence is: “We can afford to torch this much money because we will earn it back from repeat buyers — which we only will if the product is genuinely good.” Expenditure transmits quality precisely because it is wasteful from a pure-information standpoint.
Awareness vs signaling unbundling. The ad performs two distinct jobs: awareness (consumers learn the product exists) and signaling (the burn reveals confidence in quality). The dissipative-signaling result concerns the signaling job only; the awareness job has informational value even with no asymmetry to resolve. This unbundling becomes load-bearing under third-party verification (see alternative structures).
Stability. Profitable deviations: none for either type within the separation band — H’s +3 beats its No-Ad 0, L’s No-Ad 0 beats its Ad −2. The only deviations that bite are off-band: if C falls below π′, L’s one-shot fooling deviation becomes profitable and separation breaks; if C exceeds 2π, even H walks away. Off-path, the Intuitive Criterion pins beliefs so the high type cannot be deterred from the least-cost separating burn.
Reader-reproducibility check. A reader can reconstruct this equilibrium from the components above: the four payoff orderings (V_L < C ≤ V_H; s_L < 0 < s_H), the prior p relative to the threshold p* = −s_L/(s_H − s_L), the PBE consistency conditions, and the Cho–Kreps off-path refinement together pin both the surviving separating equilibrium and the defeat of both pooling branches.
Probability discipline. Probabilities appear at exactly one node: Nature’s type draw, Pr(H) = p, Pr(L) = 1−p (chance node). Every other edge is a decision — the firm’s burn-level choice e (including the binary Ad/No-ad) and the consumers’ Buy/Don’t are choices and carry no probabilities. The buy threshold p* and the posteriors μ are belief cutoffs the consumer computes, not probabilities stamped on decision edges. (Mixed strategies, if introduced, are strategy objects placing randomization on a player’s choice, not exogenous chance edges — and none of the equilibria require mixing.)
Confidence: high. The PBE + Cho–Kreps derivation is standard and reproducible from the stated payoffs for both firm and consumer nodes.
Bounded-rationality note: the equilibrium above assumes consumers invert Bayes’ rule over an ad budget they cannot precisely observe. They don’t. Real consumers run a fast-and-frugal heuristic — “a company that can afford the Super Bowl must be big, confident, and here to stay” — that approximates the equilibrium inference without the arithmetic. It is ecologically rational precisely because the separating equilibrium makes it true on average; it tracks the signal’s direction, not its magnitude. The heuristic failure modes a hyperrational model misses: (a) consumers respond to celebrity/affect rather than cost, letting a deep-pocketed L firm buy a misleading signal — and per the sharpened condition, this scalable L can defeat separation in the short fooling window (π′ > C), exiting before the repeat stream exposes it; (b) real ads are not pure burns — they carry affect, mere-exposure familiarity, brand association, so observed spend = signal + persuasion + salience, and you can’t read the whole budget as dissipative; (c) firm-side agency distortion — ad budgets set by CMOs with empire-building incentives and availability bias can sit outside the separation band (over-spend no quality story justifies); (d) consumers see only “expensive,” not actual cost, so the signal is coarse. None overturn the equilibrium; they widen its error bars and locate the live exploit (the scalable, well-capitalized L) the sharpened condition now names directly. The clean separation band is the rational benchmark, not a forecast — and the recommendations in the final section account for this.
Credibility assessment
- credibility: the content-free Super Bowl ad burn — credible. Commitment device or future-shadow: Schelling commitment via sunk, non-recoverable expenditure combined with differential benefit — the low type physically cannot recoup C from one-shot sales. The commitment is the money already on fire; mimicry is observable-and-unprofitable, which is exactly the credibility test — provided C clears the sharpened π′ bar.
- credibility: the verbal claim “our product is high quality” — cheap talk. Why dismissible: no commitment device, no sunk cost, no future-shadow. Both types say it at zero extra cost, so it moves no belief; no cost asymmetry ⇒ no separation.
- credibility: “We stand behind this for years” / unbacked guarantee — cheap talk unless enforced. Why dismissible as stated: an announcement without a commitment device. It becomes credible only when bonded — a third party (courts, FTC) makes reneging costly, or a refund converts it into a costly action.
- credibility: a money-back guarantee / warranty — credible, and often superior to the ad. Commitment device: a warranty is a cost that bites the low type differentially (more returns), satisfying single-crossing in a contractible form. Warranty equilibrium sketch: a warranty of depth d triggers return with probability r_θ where r_L > r_H; expected refund burden r_L·d > r_H·d at every depth, so a depth in the band where r_L·d > (gain from fooling the buyer) ≥ r_H·d deters the low type while the high type bears it comfortably. Crucial difference: the warranty is a transfer (refund money moves to dissatisfied buyers, rarely triggered when quality is genuinely high) rather than a pure burn (the ad money is gone either way) — same separating architecture, strictly lower deadweight loss when quality is high.
Schelling’s distinction: an announcement is not a threat or promise until something makes non-performance costly. The Super Bowl ad is a promise made credible by being prepaid and unrecoverable; “trust us” is a promise made of air.
Why cheap talk fails, formally. A costless message has identical (zero) cost for H and L, so no single-crossing holds; both types send the most favorable message → posterior = prior → uninformative babbling equilibrium. By Crawford–Sobel (1982, Econometrica 50(6):1431–51), cheap talk transmits information only when sender and receiver interests are sufficiently aligned — informativeness shrinks as sender bias grows and collapses to babbling once bias is large enough. Here interests conflict on the low type (every firm wants consumers to buy; L actively wants to be mistaken for H), so talk fully unravels. The only way to restore information is to attach a cost the low type won’t pay — make the talk not cheap. That is the formal sense in which burning money says more than words.
Confidence: high on all credibility labels; citations web-confirmed.
Alternative structures
- Alternative classification — duration, genuinely one-shot (no repeat purchase). What changes: remove period 2 (a search good, or a product nobody rebuys) and repeat business vanishes, so V_H = V_L (both types earn only π from a fooled buyer). Single-crossing collapses — the band V_L < C ≤ V_H degenerates to a knife-edge; for any C below it both types advertise (pooling), above it neither does. With μ(Ad) = μ(No Ad) = p and symmetric payoffs, no type strictly prefers Ad, so the informative separating PBE does not exist — advertising cannot signal quality. Implication for the dominant analysis: the dominant equilibrium is contingent, not robust — the repeat-purchase assumption is doing all the work. This is a testable prediction matching the Nelson / Milgrom–Roberts / FTC literature: dissipative advertising signals quality for experience goods with repeat purchase (beer, soda, cars, insurance) and essentially not for one-shot / search goods verified by inspection.
- Alternative classification — information structure, verifiable quality / cost-side single-crossing. What changes: (a) if a credible third-party channel (Consumer Reports, verified reviews, certification) lets consumers observe θ directly, the asymmetry disappears and μ is set by the channel. Here the awareness/signaling unbundling becomes load-bearing: under verification the correct statement is not “the whole ad spend is deadweight” but “the signaling premium in the spend is deadweight” — H rationally shifts from an expensive burn down to a cheaper awareness-only buy, because the channel now does the separating. Review aggregators are thus an equilibrium-shifting (not merely equilibrium-destroying) entrant — they strip the dissipative premium while leaving the informational/awareness floor. (b) Alternatively flip to genuine cost-side single-crossing (production cost of high quality actually higher): the signal can be price or production lavishness, yielding a Spence-style separating equilibrium even without repeat purchase. Implication for the dominant analysis: different mechanism, same equilibrium architecture, telling you which signal to reach for.
- Alternative classification — duration, infinite horizon (static-vs-repeated framing, Axelrod / Kreps–Wilson). What changes: extend to infinite horizon and reputation substitutes for the one-time burn — an established H firm sustains demand through accumulated reputation and needs less dissipative spend, while a new entrant with no reputation must burn more to separate. Implication for the dominant analysis: prediction is that Super Bowl advertising should skew toward new entrants, repositionings, launches and toward incumbents defending against entry — not toward mature brands with secure reputations, which can pool-on-No-Ad profitably. The naive static one-shot reading would wrongly predict no advertising at all.
Strategic recommendations
- Calibrate spend into the (π′, 2π] window — mechanism it leverages: commitment / sunk-cost lever. The ad must visibly exceed the largest one-shot fooling revenue a low-quality imitator could grab (π′), not merely your nominal one-period profit, to deter imitation — but not exceed your two-period return 2π. Confirm C ≤ V_H at your own repeat rate before committing; if lifetime value doesn’t clear the slot price, the signal loses money even when you’re telling the truth. Expected equilibrium shift: underspending pools you with L; overspending destroys your surplus. The expense is the point — don’t seek a “cheaper equivalent,” cheapness destroys the cost asymmetry. (Watch the slot seller — it will try to price you toward 2π and capture the rent.)
- Spend only on repeat-purchase products — mechanism it leverages: duration lever. For one-shot / search goods the strategy has no separating power; the burn signals nothing. Expected equilibrium shift: redirect the budget. For one-shot products, switch the mechanism to a money-back guarantee / warranty, which restores single-crossing on the return rate (r_L > r_H) and, unlike the ad, isn’t wasted when quality is high — better signal, lower deadweight loss.
- Make the burn visible and unfakeable — mechanism it leverages: credibility lever. What separates is the observable, irreversible commitment, not the creative. Expected equilibrium shift: buy the conspicuous, verifiably-expensive placement; ad content can be empty.
- Don’t substitute cheap talk — mechanism it leverages: credibility lever. Any costless claim, tagline, or unenforced guarantee lands in the babbling equilibrium (Crawford–Sobel conflict). Expected equilibrium shift: if you must communicate in words, bond them — attach a refund, penalty, escrow, or sunk cost (legally binding warranty) — so the message exits the cheap-talk class and imports a commitment cost the imitator can’t bear.
- Pre-empt the verification entrant — mechanism it leverages: information-structure lever. The separating equilibrium’s dissipative premium is fragile to cheap third-party verification. Expected equilibrium shift: if reviews/certification are becoming credible in your category, shift budget out of the burn premium into winning on the verifiable channel plus an awareness-only buy — the burn loses its signaling job once quality is observable, but you still need reach.
- As a regulator — mechanism it leverages: classification-dimension alteration (information structure via external penalty). Strengthening false-advertising penalties is a substitute for dissipative advertising — it makes cheap talk credible and lets firms stop burning the money. Expected equilibrium shift: converting cheap talk into costly talk via an external penalty rather than a private bonfire collapses the need for the burn.
- As a consumer / buyer — mechanism it leverages: equilibrium-contingent inference. Trust dissipative ad signals for repeat-purchase experience goods; heavily discount them for one-shot purchases, where the separating equilibrium doesn’t hold and the spend is uninformative.
- If you already hold reputation, consider pooling-on-No-Ad — mechanism it leverages: duration lever (outside option of standing reputation). An incumbent whose prior p is high enough that consumers buy anyway (optimistic-prior case) gains nothing from the burn. Expected equilibrium shift: advertise to defend against entrants or to launch, not to re-signal quality the market already believes.
The slot-price anchor and a surfaced tension
The same-price-for-everyone invariant is load-bearing; the dollar figure is not. One posted price for all buyers regardless of type holds at every vintage. The specific figure drifts upward and will date: ~$7M for a 30-second slot at Super Bowl LVIII (2024); ~$8M average (2025). Surfaced tension (minor, non-substantive): the 2026 figure is given as ~$8–9M in one derivation and ~$10M in the other; both treat it as illustrative and rising, and the window/band argument is vintage-independent because parameters scale together. A reader in a later year should substitute the then-current rate card. Confidence: high on the invariant; the exact current-year figure is the only moving target.
One-line synthesis. A content-free Super Bowl ad is a credible quality signal because it is a prepaid, unrecoverable burn that only a firm anticipating repeat business can recoup — placing the cost in the separation band V_L < C ≤ V_H (sharpened: π′ < C ≤ 2π) and sustaining a Perfect Bayesian separating equilibrium, the unique survivor of the Cho–Kreps Intuitive Criterion against both pooling branches in the lemons regime, where words (cheap talk) cannot separate because lying is free and both types do it. The signal works on the benefit side, depends entirely on the product market being repeated, and is substitutable by warranties (transfer, not burn) or regulation (cheap talk made costly).
References. Spence (1973, QJE 87(3):355–74, job-market signaling); Nelson (1974, JPE 82(4):729–54, “Advertising as Information,” experience-goods/repeat-purchase argument); Milgrom & Roberts (1986, JPE 94(4), price/advertising signals — Intuitive Criterion “rules out all pooling equilibria” under single-crossing); Cho & Kreps (1987, QJE, Intuitive Criterion); Crawford & Sobel (1982, Econometrica 50(6):1431–51, strategic information transmission / cheap talk). All web-confirmed in claim resolution.
(visual rendered — see artifact)