C 认知发展课程A Cognitive Development Curriculum

第 11 章 · Chapter 11

想他们在想什么

Game Theory — Thinking About What Others Think

当结果取决于同样在算计你的人,理性就不是孤立地找最优解。纳什均衡是定理,却只是规范基准而非人类行为的预测:真人只推演一两层,会拒绝不公平的报价,会自掏腰包惩罚搭便车者。

7,272 词 · 约 33 分钟 · 15 节

#1. Executive Summary

The central thesis of this chapter is that rational action, when your outcome depends on others who are also reasoning about you, is not a matter of finding the “best” move in isolation — it is a matter of forming beliefs about others’ beliefs, and choosing a strategy that survives their strategic response. Game theory is the discipline that makes this reasoning explicit. It gives us a precise vocabulary (players, strategies, payoffs, information, equilibrium) and a set of durable mental patterns (the prisoner’s dilemma, coordination, chicken, signaling) for recognizing the strategic structure of a situation.

The most important conclusions are these. First, the Nash equilibrium — the concept for which John Nash won the 1994 Nobel Memorial Prize — is a well-established mathematical result (every finite game has at least one) but only a normative benchmark for human behavior, not a reliable prediction of it. Real people, in hundreds of laboratory experiments, systematically deviate from equilibrium play: they reject unfair offers, cooperate in one-shot dilemmas, and reason only one or two steps deep. Second, the same repetition, reputation, and communication that game theory shows can sustain cooperation also explain why cooperation collapses when interactions become one-shot, anonymous, or crowded. Third, the theory’s greatest practical triumphs — spectrum auctions, kidney exchange, school matching — come from reversing the question: instead of predicting play, designers engineer the rules so that self-interested play produces good outcomes (mechanism design). Fourth, game theory is a thinking tool, not a crystal ball. Its highest value for an individual is diagnostic: it tells you what kind of game you are in, whether your threats are credible, whether a winner’s curse is hiding in your estimate, and when to play robustly rather than “optimally.”

This chapter closes Tier 1 by supplying its missing dimension: the strategic rationality of others. The Incentives chapter took the rule-author’s view; this chapter takes the player’s view inside a rule-governed world — and shows how the incentives you design evolve, under others’ strategic responses, into games you never anticipated.

#2. Why This Topic Matters

Almost everything you care about — markets, salaries, relationships, politics, war, the behavior of AI systems — is a form of strategic interaction. And in strategic settings, the single most expensive cognitive error is to model the other side as passive, fixed, or predictable. When you set a price, competitors re-price. When a regulator writes a rule, the regulated party finds the loophole (the Goodhart and cobra effects of the Second-order Thinking chapter are strategic best responses). When you make a threat, the other side asks whether you would actually carry it out. Treating other minds as furniture — as part of a static environment rather than as agents optimizing against you — is the characteristic failure that game theory is designed to correct.

This chapter completes a four-part account of the social world that the curriculum has been building:

  • Cognition (individual): how a single mind thinks and errs (Chapters 1–4, 7).
  • Rules (incentives): how designed consequences shape behavior (Chapter 8).
  • Structure (systems): how stocks, flows, and feedback generate macro-behavior (Chapters 9–10).
  • Strategic reason (others): how rational individuals choose within those structures, anticipating each other — this chapter.

The interface with Complex Systems (Chapter 10) is complementary, not competitive: complex systems explains how macro-structure emerges from many interacting parts; game theory explains how rational individual choice forms within that structure. When many heterogeneous, boundedly-rational players interact, the two merge — real strategic systems often look more like complex adaptive systems than like the clean static equilibria of the textbook.

#3. Foundations

#The core vocabulary

A game is any situation with (1) two or more players, (2) strategies (complete plans of action) available to each, (3) payoffs that depend on the combination of everyone’s choices, and (4) an information structure specifying who knows what. Two distinctions about information matter throughout:

  • Complete vs. incomplete information: do players know each other’s payoffs? (Incomplete-information games, where you are uncertain about the other’s “type,” were formalized by John Harsanyi, 1967–68, who showed how to model them as Bayesian games — applied Bayesianism, connecting directly to Chapter 4.)
  • Perfect vs. imperfect information: when it is your turn, do you know everything that has happened so far? (Chess has perfect information; poker does not.)

Games are represented in normal form (a payoff matrix, best for simultaneous moves) or extensive form (a game tree, best for sequential moves).

Dominance is the first tool: a strategy is dominant if it is your best choice no matter what others do. If you have one, use it. Most interesting games have none.

The Nash equilibrium (Nash, 1950) is the central solution concept: a combination of strategies such that no player can do better by unilaterally changing their own strategy, holding others’ fixed. It is a mutual best-response — a point of no unilateral improvement. Nash’s theorem is a well-established mathematical result: every finite game (finitely many players, finitely many strategies) has at least one equilibrium, possibly in mixed strategies. The proof uses Brouwer’s (and, in David Gale’s simplification, Kakutani’s) fixed-point theorem.

Pure vs. mixed strategies: a pure strategy is a definite choice; a mixed strategy is a probability distribution over choices. Randomizing can be genuinely rational — in matching pennies, penalty kicks, and poker, being predictable is fatal, so the equilibrium requires you to randomize in specific proportions. (This connects to Chapter 3: a mixed strategy is a probability distribution, and playing one well requires calibration.) Von Neumann’s minimax theorem (1928) established the existence of such optimal randomized play for two-player zero-sum games; Nash generalized it to any number of players and non-zero-sum games.

For sequential games, backward induction (“think from the end”) solves the tree by reasoning from the final moves backward. The resulting refinement, the subgame-perfect equilibrium (Selten, 1965), rules out equilibria supported by incredible threats — threats a rational player would not actually carry out when the moment came.

A repeated game is the same stage-game played many times. The folk theorem (a family of results developed through the 1970s–80s) captures the “repeated-game revolution”: when players are patient and the game is repeated indefinitely, a huge range of outcomes — including mutual cooperation in the prisoner’s dilemma — can be sustained as equilibria, enforced by the threat of future punishment. This is a well-established theorem, but note its double edge: it says cooperation can be sustained, and also that almost anything else can be too. It is a possibility result, not a prediction.

Other foundational terms: commitment (removing your own options to change others’ expectations), credibility (whether others believe you will do what you say), signaling (using costly actions to communicate private information), mechanism design (designing the rules so equilibrium play yields a desired outcome), the winner’s curse (systematic overpayment in common-value auctions), and the evolutionarily stable strategy or ESS (a strategy that, once common, cannot be invaded by a rare mutant).

#Historical development

  • 1838 — Cournot’s duopoly. Antoine Augustin Cournot analyzed two firms choosing quantities, anticipating the equilibrium concept by more than a century.
  • 1928 — von Neumann’s minimax theorem. John von Neumann proved that every two-player zero-sum game has a value and optimal (possibly mixed) strategies.
  • 1944 — Theory of Games and Economic Behavior. Von Neumann and Oskar Morgenstern founded the field as a discipline, introducing expected-utility theory and cooperative game theory.
  • 1950 — Nash equilibrium. John Nash’s two-page PNAS paper “Equilibrium Points in N-Person Games” defined and proved existence of the equilibrium that bears his name. The “Nash program” sought to ground cooperative outcomes in non-cooperative foundations.
  • 1960 — the Cold War turn. Thomas Schelling’s The Strategy of Conflict reframed game theory around commitment, credible threats, focal points, and deterrence — insights drawn from and applied to nuclear strategy.
  • 1960s–70s — refinements. Reinhard Selten (subgame perfection, 1965; trembling-hand perfection) and John Harsanyi (incomplete-information/Bayesian games, 1967–68) tightened the theory. Nash, Selten, and Harsanyi shared the 1994 Nobel.
  • The applied-economics revolution. Auction theory and market design (Paul Milgrom, Robert Wilson, Roger Myerson, Alvin Roth, Lloyd Shapley) turned theory into working institutions; Milgrom and Wilson won the 2020 Nobel, Roth and Shapley the 2012.
  • The behavioral turn. Colin Camerer’s Behavioral Game Theory (2003) synthesized decades of experiments showing how real people play.
  • The computer-science merger. Algorithmic game theory, mechanism design for online markets, and superhuman poker AIs (Libratus 2017, Pluribus 2019) mark the modern frontier.

#The canonical games as payoff matrices

Prisoner’s dilemma. Two suspects each choose Cooperate (stay silent) or Defect (confess). Numbers are payoffs (higher is better); the first entry is the Row player’s, the second the Column player’s.

Column: CooperateColumn: Defect
Row: Cooperate3, 30, 5
Row: Defect5, 01, 1

Defect is dominant for both (5 > 3 if the other cooperates; 1 > 0 if the other defects), so the unique equilibrium is (Defect, Defect) = (1, 1) — worse for both than (Cooperate, Cooperate) = (3, 3). Individual rationality produces collective disaster. This is the structure of arms races, price wars, and overfishing.

Tragedy of the commons is the N-player prisoner’s dilemma: each herder gains from adding one more animal, but the shared pasture collapses when all do. This is where the Systems Thinking chapter’s commons archetype acquires its strategic form.

Chicken / brinkmanship. Two drivers speed toward each other; each chooses Swerve or Straight.

Column: SwerveColumn: Straight
Row: Swerve0, 0−1, +1
Row: Straight+1, −1−10, −10

There are two pure equilibria — (Straight, Swerve) and (Swerve, Straight) — and mutual destruction if both go straight. The strategic prize goes to whoever can credibly commit to going straight (e.g., by conspicuously ripping out the steering wheel). This is the nuclear-standoff structure.

Stag hunt (coordination). Both hunters do best by jointly hunting stag, but hunting stag alone yields nothing; hare is a safe solo option.

Column: StagColumn: Hare
Row: Stag4, 40, 3
Row: Hare3, 03, 3

Both (Stag, Stag) and (Hare, Hare) are equilibria. (Stag, Stag) is better for everyone but riskier; (Hare, Hare) is safe. Which one a society lands on — the standards war, the language, the convention — is a matter of trust and focal points, not calculation.

Ultimatum game. A proposer offers a split of, say, $10; the responder accepts (both paid as proposed) or rejects (both get $0). The self-interest equilibrium is “offer the smallest positive amount; accept anything positive.” Real players violate this systematically — the door into behavioral game theory (Section 9.2).

#4. Current Scientific Understanding

It is essential to separate what is securely known from what is plausible and what is contested — and to apply the chapter’s twofold tiering discipline: mathematical theorems are secure about their models, but this lends no automatic credibility to claims about real human behavior.

The well-established core:

  • Equilibrium existence and the folk theorem are proven theorems (secure about models).
  • Auction empirical regularities: the winner’s curse is real and robust in laboratory common-value auctions (Kagel & Levin and successors); revenue equivalence holds under its assumptions and breaks predictably when they fail.
  • Cross-cultural ultimatum findings: the pattern that proposers offer substantial shares and responders reject low offers is one of the most replicated results in experimental economics — in Western (WEIRD) samples, mean offers cluster around 40% and offers below ~20% are rejected roughly half the time — but its magnitude varies enormously across societies (Henrich et al., 2001, 2005).
  • Evolution-of-cooperation results: the mathematical conditions under which cooperation can evolve (Maynard Smith, Axelrod & Hamilton 1981, Nowak 2006) are well-established as theory.

The plausible (supported but not settled):

  • Mechanism design’s real-world successes — spectrum auctions, kidney exchange, school choice — are genuine, but their success depended heavily on context and careful engineering, and some celebrated designs failed when transplanted.
  • Levels-k / cognitive hierarchy as a description of bounded strategic reasoning is well-supported in guessing games but is a modeling framework, not a law.

The contested:

  • Whether equilibrium predicts behavior in the wild. In many field settings the assumptions are untestable and the predictions unfalsifiable.
  • Whether behavioral deviations are “mistakes” or alternative rationalities. Rejecting an unfair offer is irrational only if you assume the player values nothing but money — a substantive, and probably false, assumption.
  • The lab-to-field gap. Effects measured on undergraduates in one-shot lab games may not generalize; the Bardsley/List critique of dictator-game “altruism” (below) is a cautionary case.

#5. Interdisciplinary Perspectives

Four disciplines study strategic interaction, and they fit together like grammar, field, deviation, and deep history.

  • Mathematics — the grammar. Provides the equilibrium concepts, existence proofs, and the logic of dominance, backward induction, and repeated games. It tells you what is coherent, not what is true of people.
  • Economics — the field. Tests and applies the grammar in markets: auctions (spectrum, treasury, advertising), oligopoly pricing, labor markets, and matching. It is where mechanism design turns theory into institutions.
  • Psychology & cognitive science — the deviations. Documents how real humans depart from equilibrium: ultimatum rejections, one-shot cooperation, bounded (level-1/2) reasoning, and the role of emotion and fairness. Camerer, Fehr, Gächter, and Henrich are central.
  • Evolutionary biology — the deep history. Asks why we have the preferences we do: ESS analysis, hawk-dove, and the mechanisms by which cooperation and even fairness intuitions could have been selected for.

Where they complement: biology explains the origin of the fairness preferences that psychology measures and that economics must build into its models if mechanism design is to work on real humans. Where they conflict: the mathematician’s clean equilibrium vs. the psychologist’s messy, heterogeneous player; the economist’s revealed preferences vs. the biologist’s evolved functions; the designer’s optimization vs. a population of emotional, norm-driven agents who may punish, spite, and cooperate against their own material interest.

#6. Mental Models

Each model below comes with when it works and where it fails.

The game-classifier. Before analyzing, classify: is this coordination (we both want to align — find the focal point), conflict/zero-sum (my gain is your loss — find your BATNA and exit), or mixed-motive (elements of both — find the cooperation zone)? Works: almost always, as a first orienting step. Fails: when you misread a positive-sum game as zero-sum (a very common and costly error in negotiation).

Nash equilibrium as the “no unilateral improvement” test. Ask: given what everyone else is doing, would anyone want to change? If not, you have found a stable point. Works: identifying stable configurations and traps. Fails: when there are many equilibria (it cannot tell you which one), and as a prediction of what people will actually do.

Backward induction (“think from the end”). In sequential situations, reason from the last move backward. Works: finite games with clear endpoints (negotiations with deadlines, litigation). Fails: when players are not perfectly rational (the centipede game and finitely-repeated prisoner’s dilemma show humans routinely violate it) and when the end is unknown.

Commitment and burning bridges (Schelling). Reducing your own options can strengthen your position by changing others’ expectations. Works: when the commitment is visible and irreversible. Fails: when it is not credible, or when it removes flexibility you later need.

The shadow of the future. The prospect of continued interaction disciplines present behavior. Works: ongoing relationships, repeat business. Fails: known endpoints, anonymity, one-shot encounters.

Signaling and costly signals. Talk is cheap; to communicate credibly, use actions costly enough that a bluffer would not imitate them. Works: education as a signal, warranties, luxury goods, sincere apologies (costly ones). Fails: when the “signal” is cheap enough to fake.

Winner’s-curse awareness. In common-value competition, winning means your estimate was the highest — probably too high. Shade your bid. Works: auctions, M&A, competitive hiring, bidding wars. Fails: if you forget it — overconfidence (Chapter 2) makes this the default error.

BATNA and reservation price. Your power in any negotiation is your best alternative to a deal. Works: every negotiation. Fails: when you have not actually developed your alternative, only imagined it.

The focal point (Schelling). Sometimes an equilibrium is found, not calculated — a solution salient enough that everyone expects everyone to expect it. Works: coordination without communication (meeting-place problems, market conventions). Fails: when salience differs across cultures or people.

“Play robustly, not optimally, when the model is uncertain.” When you cannot fully model the game, choose strategies that do acceptably across many scenarios rather than optimize for one. Works: deep uncertainty (connect to Chapter 1’s grades of uncertainty). Fails: rarely — this is the humility rule, and its main risk is being used as an excuse to avoid analysis.

#7. Common Misconceptions

“Game theory predicts what people will do.” No. Its core equilibrium concept is a normative benchmark. It predicts well in some structured, high-stakes, repeated settings (professional sports serves, some auctions) and poorly in many others. Intelligent people fall for this because the mathematics is so elegant that its prestige seems to transfer to its behavioral claims — the exact error the tiering discipline is designed to prevent.

“Game theory assumes people are selfish.” No. It assumes people have preferences and pursue them consistently. Fairness, spite, love, and revenge can all be written into the payoffs. The theory is content-free about what people value; the “selfishness” is an assumption some modelers add, not a feature of the framework.

“Nash equilibrium is the best outcome.” No — it is the stable outcome, often Pareto-inferior. The prisoner’s dilemma’s equilibrium (both defect) is worse for both players than mutual cooperation. Stability and optimality are different properties.

“Game theory is just for economists and mathematicians.” No — its most valuable use is everyday diagnosis: recognizing the structure of a conflict or a coordination problem.

“All games are zero-sum.” No — most of life is mixed-motive or positive-sum. Treating a negotiation as pure conflict destroys value both sides could have shared.

“Rationality means winning at all costs.” No — rationality means choosing consistently to achieve your goals, which for real humans include relationships and reputations.

“Game theory is amoral.” The framework is value-neutral; the user is not. The same analysis that helps a predator can help a cooperator build trust. Good game theory models trust and reciprocity as readily as exploitation.

#8. Real-World Applications

Negotiation and conflict. Know your BATNA and estimate theirs; anchor with informed first offers (Chapter 2); use credible commitments; be willing to walk away. Reframe zero-sum into positive-sum where value can be created.

Cooperation problems. Teams, commons, and climate are N-player prisoner’s dilemmas. Solutions come from lengthening the shadow of the future, adding monitoring, enabling reciprocity and (careful) punishment, and shrinking the group so reputation bites.

Markets and pricing. Oligopolists face a repeated prisoner’s dilemma; price wars are defection cascades. Know when to bid and when to walk, and always ask whether a winner’s curse lurks in your valuation.

Careers and signaling. Credentials, reputations, and visible effort are costly signals. Their value lies precisely in being hard to fake.

Relationships. Trust games and coordination problems abound; the shadow of the future (expecting to interact again) is what makes trust rational.

Online platforms. Reputation systems (eBay feedback, reviews, ratings) are engineered shadows of the future — mechanisms that make one-shot encounters behave like repeated ones.

AI. Multi-agent systems negotiate, price, and bluff; alignment can be framed as a mechanism-design problem (designing objectives so that a powerful optimizer’s best response is what we want). Algorithmic pricing raises the novel question of machine collusion.

Public policy. Market design now allocates radio spectrum, medical residencies, school seats, and organs (Alvin Roth and Lloyd Shapley’s Nobel-winning work on the Gale–Shapley deferred-acceptance algorithm redesigned New York City’s high-school match in 2003, cutting the number of students assigned to schools they had not ranked by about 90%), and underpins proposals for carbon markets.

#9. Case Studies

#9.1 Success of the science — and its fragility: the European 3G spectrum auctions (2000)

The UK’s third-generation (3G) mobile-spectrum auction, which concluded on 27 April 2000, was designed with heavy input from economists including Ken Binmore and Paul Klemperer. It raised £22.5 billion ($34 billion, or 2.5% of GNP) — as Binmore & Klemperer put it in “The Biggest Auction Ever: The Sale of the British 3G Telecom Licences” (Economic Journal, 2002), the sum “was widely described at the time as the biggest auction ever.” Germany’s UMTS auction, months later, raised the even larger sum of €50.8 billion (99.37 billion Deutsche Mark) — as Grimm, Riedel & Wolfstetter noted, “the German finance minister cashed in the record sum of €50.8 billion.”

The instructive part is the contrast. Klemperer’s analysis (“How (Not) to Run Auctions: The European 3G Telecom Auctions,” 2002) estimated that market valuations implied revenues “could probably have been in the range 400–650 Euros per capita… in all these countries,” yet realized revenues ranged from more than €600 per capita (UK, Germany) down to roughly €20 per capita in Switzerland — a difference of more than an order of magnitude for licenses of broadly similar value. Switzerland was “a complete flop… where only four bidders showed up to bid for four licenses” (Grimm, Riedel & Wolfstetter). The difference was auction design and sequencing, not the value of what was sold. Countries that copied the theory carelessly, or ran their auctions after market sentiment had turned and after firms had learned to coordinate, raised far less. Klemperer’s own verdict was that other governments “would be foolish not to copy the U.K. in auctioning the radiospectrum, but they would be equally foolish to blindly copy the U.K. design without attention to their local circumstances.”

The lesson is twofold. Mechanism design is powerful — good rules extracted enormous value and allocated spectrum efficiently — and fragile — the same theory, applied without attention to entry, collusion, and timing, failed. There is also a second-order twist: several telecom firms that “won” and paid the highest prices arguably suffered a winner’s curse, overpaying for licenses whose value collapsed when the telecom bubble burst. The auction succeeded for the sellers (governments) precisely because it may have cursed the winners.

#9.2 Success as natural experiment: the ultimatum game across small-scale societies (Henrich et al. 2001, 2005)

In the ultimatum game (introduced by Güth, Schmittberger & Schwarze, 1982), a proposer offers a split of a sum; the responder accepts (both are paid) or rejects (both get nothing). The self-interest prediction is stark: the responder should accept any positive offer, so the proposer should offer the minimum. In Western samples, this prediction fails: mean offers cluster around 40–50%, and low offers (under ~20%) are rejected roughly half the time.

Joseph Henrich and colleagues took the game to 15 small-scale societies on four continents. The results dethroned two opposite dogmas at once. Against “humans are selfish,” no society played the self-interest equilibrium. Against “humans are uniformly fair,” behavior varied enormously and tracked culture. The Machiguenga of Peru offered a mean of about 26% with almost no rejections (roughly 5% rejected). Among some Melanesian groups (the Au and Gnau), proposers made hyper-fair offers exceeding 50% — and these were sometimes rejected, because accepting a large gift carried an obligation of subordination. Among the whale-hunting Lamalera of Indonesia, whose livelihood requires massive cooperation, 63% of proposers divided the pie equally and the mean offer was about 57%. The best predictors of fairness were a society’s degree of market integration and payoffs to cooperation in daily life. Fairness, in other words, is real, universal in existence, and profoundly variable in form — a cultural product, not a fixed human constant.

#9.3 Failure of the “prediction culture”: a real winner’s curse

The winner’s curse was named by three petroleum engineers — Edward Capen, Robert Clapp, and William Campbell — in a 1971 paper, “Competitive Bidding in High-Risk Situations” (Journal of Petroleum Technology), after their firm kept overpaying for offshore oil leases. In a common-value auction, the asset is worth roughly the same to everyone, but each bidder has a noisy estimate. The winner is, by construction, likely the bidder with the most optimistic (highest) estimate — so winning is bad news about your estimate. Bidders who fail to shade their bids to account for this “win, then lose money, then curse.”

The corporate scale-up is the AOL–Time Warner merger. Announced on 10 January 2000 at the peak of the dot-com bubble, it was a roughly $165 billion stock merger (reported at up to ~$180 billion, at about $108 a share). When the bubble burst and the promised synergies never materialized, AOL Time Warner took a goodwill write-down of about $99 billion and reported a net loss of $98.7 billion for fiscal year 2002 — the largest annual loss any company had ever reported, a record that still stands. The mechanics are pure common-value overpayment amplified by overconfidence (Chapter 2) and escalation: a euphoric estimate of a hard-to-value asset, a competitive/visionary impulse to close the deal, and no adequate discount for the possibility that the very enthusiasm driving the price was a warning sign.

#9.4 Failure of equilibrium thinking: brinkmanship and the Cuban Missile Crisis

The 1962 Cuban Missile Crisis is the canonical illustration of chicken/brinkmanship — and of the limits of clean equilibrium reasoning. Schelling analyzed such standoffs with his concept of “the threat that leaves something to chance”: when a direct threat (“I will launch a nuclear war”) is not credible because carrying it out would be suicidal, an actor can instead manipulate risk — take steps (like a naval blockade) that raise the probability of an accidental, uncontrolled escalation, making the shared danger itself the coercive instrument.

The honest caveat is crucial. Schelling read the outcome as a US “win” achieved through adroit risk-manipulation, but later scholars (e.g., the game-theoretic reappraisals collected in A Game-Theoretic History of the Cuban Missile Crisis, Zagare) have argued that this reading is both theoretically strained and empirically thin — there is scant evidence the Kennedy administration calibrated risk with anything like “mathematical precision,” and the crisis involved genuine loss of control (a U-2 was shot down; a Soviet submarine nearly launched a nuclear torpedo). We know the outcome; the players did not. This is the deepest lesson of the case: equilibrium analysis of brinkmanship describes a logic, but real crises are run by frightened humans with imperfect information and imperfect control — closer to a complex system near a tipping point (Chapter 10) than to a solved game tree.

A quieter, everyday version is the airline price war: a repeated prisoner’s dilemma in which one carrier cuts fares, rivals match, and all bleed money in an outcome none of them rationally wanted but none could unilaterally exit.

#9.5 Personal scale: tit-for-tat with a colleague

Consider a recurring interaction — say, covering shifts, sharing credit, or reciprocating favors with a coworker. This is an iterated prisoner’s dilemma. Robert Axelrod’s computer tournaments (first run in 1980, published in The Evolution of Cooperation, 1984) found that the simple tit-for-tat strategy — submitted by Anatol Rapoport, it cooperates first, then copies the other’s last move — won both times. Its virtues: it is nice (never defects first), retaliatory (punishes defection immediately), forgiving (returns to cooperation the moment the other does), and clear (easy to read, so the other learns to cooperate). Applied to your colleague: open cooperatively, respond proportionally to defection, forgive quickly, and be legible about your rule. (Caveat below: tit-for-tat’s dominance is contingent, not universal.)

#10. Practical Framework: The Strategic Reading Protocol

A seven-step routine for reading any strategic situation before you act.

Step 1 — Name the game. List all players, including hidden ones (the boss behind the negotiator, the future rivals watching this precedent). For each, ask what they actually want and what options they have. Where possible, sketch the payoff table.

  • Reflective questions: Who else is affected by or watching this? What does each party truly value (not what they say)? What are the actual options on each side?

Step 2 — Classify it. Coordination (→ find the focal point), conflict (→ find the BATNA and the exit), or mixed (→ find the cooperation zone)? Zero-sum or positive-sum?

  • Reflective questions: Is there value to be created, or only divided? Where do our interests secretly align? Am I mistaking a positive-sum game for a fight?

Step 3 — Run the best-response check. For each of your options, what is their best response? Then go one level deeper — but stop at level 3. Real people reason only one or two steps deep (see levels-k below); assuming infinite rationality is itself a modeling error.

  • Reflective questions: What is their best reply to my move? What do they think I will do? Am I over-thinking a shallow opponent (or under-estimating a deep one)?

Step 4 — Ask the credibility question. Which of your threats and promises would you actually carry out? Which of theirs would they? Discard the incredible ones — or make them credible with a commitment device.

  • Reflective questions: Would I really do this if called? What makes my commitment believable? Can I remove my own escape route to strengthen my position?

Step 5 — Extend the shadow of the future. Can you make the interaction repeated, observable, and reputation-bearing (which fosters cooperation)? Or, if you are being exploited in a one-shot, can you end the game?

  • Reflective questions: Will we meet again, and does the other side know it? Who will hear about how I behaved? Should I be lengthening this relationship or exiting it?

Step 6 — Check information asymmetries. Who knows what? Does the structure reward honesty or bluffing? Is there a winner’s curse hiding in your estimate?

  • Reflective questions: What do they know that I don’t? If I “win,” what will that have told me about my own valuation? Am I the sucker at the table?

Step 7 — Decide. If the model is clear, play the equilibrium. If the model is uncertain, play robustly: diversify, keep slack, avoid irreversible commitments (connect to Chapter 1’s reversible/irreversible sorting).

  • Reflective questions: How confident am I in my model of the other players? What’s my downside if I’m wrong? Which choice keeps the most options open?

#One-page game-classifier chart

If the situation is…The key question is…The main tool is…Watch out for…
Coordination (want to align)Which equilibrium will we both pick?Focal points, communicationMultiple equilibria; a bad one can lock in
Pure conflict / zero-sumWhat’s my best alternative?BATNA, mixed strategies, walking awayAssuming the pie is fixed when it isn’t
Mixed-motive (e.g. PD)Can we sustain cooperation?Shadow of the future, reciprocity, commitmentOne-shot, anonymity, known endpoints
Common-value biddingIs my estimate the highest for a reason?Shade the bidWinner’s curse, overconfidence
Sequential / deadlineWhat happens at the end?Backward induction, credible commitmentIncredible threats; irrational opponents

#Two-week exercise

For fourteen days, model one real interaction per day — a negotiation, a group decision, an online exchange — using the seven steps. Keep a short journal entry for each: which step or framework caught something you would otherwise have missed? By day 14 you will have a personal catalogue of the games you actually play most.

#Negotiation drill

Before your next real negotiation, write down three things on one card: (1) your BATNA (your best option if this deal collapses), (2) their BATNA (your best estimate of theirs), and (3) your reservation price (the walk-away point implied by your BATNA). Do not open the negotiation until the card is filled. The concept comes from Fisher and Ury’s Getting to Yes (1981): “the party with the best BATNA is the more powerful party in the negotiation.”

#11. Criticisms and Limitations

Equilibrium multiplicity. Many games have multiple Nash equilibria, and the theory alone cannot say which will occur. Selecting among them requires focal points — which are psychology and culture, not mathematics. This is a fundamental, not a technical, limitation.

The rationality assumption’s empirical weakness. The evidence that people reason to equilibrium is weak in many settings. The p-beauty contest (Nagel, 1995, “Unraveling in Guessing Games,” American Economic Review) is the cleanest demonstration: players pick a number from 0 to 100 aiming for two-thirds of the average; the unique equilibrium is 0, but real players cluster at ~33 (level-1) and ~22 (level-2), reasoning only one or two steps. This “levels-k” / cognitive-hierarchy finding is robust and directly limits equilibrium as a prediction.

“Nash is rarely observed in the wild.” In field settings the assumptions are often untestable, and the theory can become unfalsifiable — able to rationalize any outcome after the fact by adjusting assumed payoffs or beliefs (a Chapter 7 red flag). A vivid instance of lab fragility is the dictator game: apparent “generosity” collapses when List (2007) and Bardsley (2008) simply add an option to take money from the recipient — the share of dictators transferring a positive amount falls from about 75% to about 35%, suggesting the giving was partly an artifact of a constrained choice set rather than pure altruism.

The difficulty of modeling real information. Specifying who knows what, and what they believe others believe, is often the hardest and most arbitrary part of an analysis — and small changes can flip the predicted outcome.

The ethical boundary. The framework is value-neutral; the user supplies the values. It can be used to build trust or to exploit it.

The risk of strategic paranoia. Habitually modeling every human interaction as a game to be won can corrode the very trust and reciprocity that make cooperation possible. The honest answer is that good game theory models trust, reciprocity, and fairness — but the dispositional risk to the person who over-applies it is real. Use it to understand, not to reduce every relationship to a maneuver.

The complexity boundary. From Chapter 10: in genuinely complex adaptive systems, optimal strategies are often unknowable in principle. There, robust strategies replace optimal ones — not as a compromise, but as the correct response to irreducible uncertainty.

No perspective here is the whole truth. Game theory is one powerful lens among several.

#12. Future Directions

Algorithmic game theory and multi-agent AI. The transition point is already behind us: Libratus (Brown & Sandholm, Science, 2018; competition January 2017) beat four top professionals at heads-up no-limit poker by a decisive 147 milli-big-blinds per hand with 99.98% statistical significance, and Pluribus (Brown & Sandholm, Science, 2019) did so in six-player poker — mastering bluffing and imperfect information, the features that make poker resemble real strategic life. Autonomous agents that negotiate, price, and bluff are now practical, raising the pressing question of algorithmic pricing collusion: independent pricing algorithms can, in simulation, learn to sustain supra-competitive prices without communicating (Calvano, Calzolari, Denicolò & Pastorello, American Economic Review, 2020), and Assad et al. found real German retail-gasoline margins rose by roughly a third where competitors adopted algorithmic pricing — a machine-scale folk theorem with real antitrust implications.

Mechanism design for the AI age. Framing alignment as game design — engineering objectives and environments so that a capable optimizer’s equilibrium behavior is beneficial — is a live research direction and the bridge to the AI Alignment tier of this curriculum.

Behavioral game theory meets cognitive science. Integrating levels-k reasoning with process models of cognition (attention, memory, emotion) promises better predictions of when people will and won’t reach equilibrium.

Climate cooperation. Designing the international climate game so that the shadow of the future — via monitoring, repeated interaction, and reputational stakes — sustains cooperation is among the highest-stakes applications of the folk-theorem logic.

The origins of moral psychology. Whether evolutionary game theory can fully explain the origin of human fairness intuitions, spite, and coalitional instincts remains open — and is the bridge to the Evolutionary Psychology chapter.

#13. Evolutionary Game Theory (Bridge to the Next Tier)

A separate strand deserves its own note because it reframes the whole enterprise. Evolutionary game theory drops the assumption of rational calculation and asks instead which strategies survive and spread in a population. Its central concept, the evolutionarily stable strategy (ESS) (Maynard Smith & Price, 1973, “The Logic of Animal Conflict,” Nature), is a strategy that, once common, cannot be invaded by a rare mutant. The hawk-dove game showed that “limited war” (ritualized, non-lethal conflict) can be individually rational, not just good for the species — resolving a long-standing puzzle in biology.

Why can cooperation evolve at all, given that selection seems to favor defectors? Martin Nowak’s synthesis (“Five Rules for the Evolution of Cooperation,” Science, 2006) identifies five mechanisms: kin selection (help relatives), direct reciprocity (I help you, you help me — the repeated game), indirect reciprocity (reputation — I help you because others are watching), network reciprocity (cooperators clustering in space), and group selection. Axelrod & Hamilton (1981, “The Evolution of Cooperation,” Science) had already shown, via the iterated prisoner’s dilemma, how reciprocity-based cooperation can get started, thrive, and resist invasion.

The bridge to the next chapter is this plausible-theory claim: game theory can explain how preferences themselves — not just strategies — might have been selected for. If fairness, spite, gratitude, and coalitional loyalty were adaptive in ancestral repeated games, humans might be born with strategic emotions pre-installed. This is precisely why the ultimatum-game responder rejects unfair offers “irrationally”: a taste for punishing unfairness, even at a cost, may be an evolved commitment device that made credible the threat that sustained cooperation. Fehr & Gächter’s altruistic punishment experiments (“Altruistic Punishment in Humans,” Nature, 2002) showed that people will pay to punish free-riders in one-shot public-goods games, and that cooperation flourishes when such punishment is possible and collapses when it is forbidden — with negative emotion as the proximate mechanism. But the picture is not uniformly rosy: Herrmann, Thöni & Gächter’s “Antisocial Punishment Across Societies” (Science, 2008), conducted in 16 participant pools worldwide, found that in some societies people also punish cooperators strongly enough to wipe out punishment’s cooperation-enhancing effect, with weak rule-of-law and weak civic norms predicting this antisocial pattern. Evolution built the machinery; culture programs how it is aimed. (This chapter previews Evolutionary Psychology; it does not do its job.)

Beginner.

  • The Art of Strategy (Dixit & Nalebuff) and its predecessor Thinking Strategically (Dixit & Nalebuff) — the best plain-language introductions; rich in real examples and the intuition behind commitment, credibility, and backward induction. Worth engaging because they build strategic instinct without heavy math.
  • Prisoner’s Dilemma (William Poundstone) — a superb narrative history of von Neumann, RAND, and the Cold War origins of the field; makes the ideas vivid and memorable.
  • A Beautiful Mind (Sylvia Nasar) — the biography of Nash; excellent on the human and intellectual context (read alongside, not instead of, the concepts).
  • Yale’s “Game Theory” open course by Ben Polak — widely regarded as the best free introduction available; lucid, rigorous, and free.

Intermediate.

  • Games of Strategy (Dixit, Skeath & Reiley) — a full textbook that keeps intuition central while adding formal tools; the natural next step.
  • The Strategy of Conflict (Schelling, 1960) — a genuine classic; readable, profound, and the origin of commitment, focal points, and deterrence theory. Worth engaging because its ideas about credibility are more useful in daily life than any equation.
  • MIT OpenCourseWare economics offerings for game theory and market design.

Advanced.

  • Behavioral Game Theory (Camerer, 2003) — the definitive synthesis of what experiments reveal about real play; essential for anyone who wants the honest behavioral picture.
  • Evolutionary Dynamics and SuperCooperators (Nowak) — the mathematics and the story of how cooperation evolves.
  • Santa Fe Institute lectures on evolutionary game theory and complex systems.

Landmark papers. Nash (1950), “Equilibrium Points in N-Person Games,” PNAS; von Neumann & Morgenstern (1944), Theory of Games and Economic Behavior; Schelling (1960), The Strategy of Conflict; Selten (1965) on subgame perfection; Harsanyi (1967–68) on incomplete-information games; Axelrod (1980–84) tournaments and Axelrod & Hamilton (1981), “The Evolution of Cooperation,” Science; Maynard Smith & Price (1973), “The Logic of Animal Conflict,” Nature; Henrich et al. (2005), “‘Economic Man’ in Cross-Cultural Perspective,” Behavioral and Brain Sciences; Fehr & Gächter (2002), “Altruistic Punishment in Humans,” Nature; Nagel (1995), “Unraveling in Guessing Games,” American Economic Review; Milgrom & Wilson’s auction-design work; Camerer (2003).

Influential researchers. John Nash, Thomas Schelling, John Harsanyi, Reinhard Selten, Robert Aumann, Paul Milgrom, Robert Wilson, Roger Myerson, Alvin Roth, Lloyd Shapley, Colin Camerer, Ernst Fehr, Simon Gächter, Joseph Henrich, Robert Axelrod, John Maynard Smith, Martin Nowak, and Elinor Ostrom (whose fieldwork on real commons governance is the empirical complement to the tragedy-of-the-commons model).

#15. Self-Check

Attempt these from memory before looking anything up.

  1. State the prisoner’s dilemma payoff structure from memory, and explain why “both defect” is the Nash equilibrium despite being Pareto-inferior to mutual cooperation.
  2. What makes a threat credible? Explain how Schelling turned commitment — reducing your own options — into a source of strategic advantage.
  3. Explain the winner’s curse in your own words. Why does winning a common-value auction carry bad news, and what should a rational bidder do about it?
  4. Describe the ultimatum game. What do its results (both in the West and across the 15 small-scale societies) tell us about the rationality assumption?
  5. When does repetition sustain cooperation in the prisoner’s dilemma — and what three conditions destroy it?
  6. Name the seven steps of the Strategic Reading Protocol and apply them to one real situation you currently face.
  7. What is the strongest argument that Nash equilibrium is not a prediction about real human behavior? (Hint: think of the p-beauty contest.)

Synthesis (not an answer key). If your answers were solid, they should weave together three threads. The machinery: an equilibrium is a mutual best-response — stable, not necessarily good — and it is a theorem about models, not a law about people. The transformations: repetition, reputation, communication, and commitment change which outcomes are reachable, turning one-shot traps into sustainable cooperation (or exposing why cooperation collapses when the shadow of the future disappears). The human boundary: real players are boundedly rational and socially motivated — they reason a step or two deep, care about fairness, and will pay to punish — so game theory’s honest role is to make strategic reasoning explicit and to diagnose the game you are in, not to forecast the future. If any thread was missing, revisit the corresponding section.

## Knowledge Card — Game Theory
- Core terms:
  - Nash equilibrium: a combination of strategies where no player can gain by unilaterally deviating (stable, not necessarily good).
  - Dominant strategy: a choice that is best regardless of what others do.
  - Mixed strategy: a rational randomization over choices (e.g., penalty kicks, poker).
  - Backward induction / subgame-perfect equilibrium: solving a sequential game from the end, discarding incredible threats.
  - Folk theorem: in indefinitely repeated games, patience lets cooperation (and much else) be sustained as equilibrium.
  - Winner's curse: in common-value auctions, winning means you likely overestimated — so shade your bid.
  - Signaling: communicating private information credibly via actions too costly for a bluffer to imitate.
  - ESS: a strategy that, once common, cannot be invaded by a rare mutant (evolutionary game theory).
- Core mental models:
  - The game-classifier: decide first whether you are in a coordination, conflict, or mixed-motive game — each needs a different playbook.
  - The credibility test: a threat you wouldn't execute is noise; strengthen it with a commitment device or drop it.
  - The shadow of the future: repetition, observability, and reputation are what make cooperation rational.
  - Play robustly, not optimally, when you cannot fully model the game.
- Connections to prior chapters:
  - Ch.1 Decision Making: precommitment gets its strategic rationale here; reversible-vs-irreversible sorting drives Step 7.
  - Ch.2 Cognitive Biases: overconfidence → the winner's curse; anchoring → first-offer effects.
  - Ch.3 & 4 Probability/Bayes: mixed strategies are probability distributions; signaling is applied Bayesianism (Harsanyi).
  - Ch.5 Behavioral Economics: social preferences are the substrate behavioral game theory measures.
  - Ch.6 Second-order Thinking: this chapter is the formal answer to "and then what — what will they do?"
  - Ch.8 Incentives: principal-agent problems are games; mechanism design builds incentives that survive strategic response.
  - Ch.9 & 10 Systems/Complex Systems: reputation is a stock; repeated games are feedback loops; equilibrium selection is a path-dependence/emergence problem.
- Recommended next chapter: Evolutionary Psychology — because game theory explains how our strategic emotions (fairness, spite, gratitude) could have been selected for, but not why they feel the way they do.
- One habit to keep: Before any consequential interaction, spend sixty seconds naming the game, classifying it, and writing down your BATNA and one credible commitment — think about what they think before you act.