The Broken Mirror: Designing Wargames That Surface Players' Flawed Self-Models
E. SokolovMost wargames test whether players can solve the problem in front of them. Fewer test whether players accurately understand themselves as decision-makers. That gap is where some of the most consequential failures live.
Photo by Polina Tankilevitch on Pexels.
Call it the broken mirror problem. Players walk into an exercise carrying a self-model: a belief about how they reason under pressure, what they prioritize, how much ambiguity they can tolerate before they anchor on bad information. That self-model is almost always wrong in predictable ways. Designers who ignore it are leaving one of the richest failure surfaces untouched.
This is a design problem, and it has a tractable solution set.
Why the Self-Model Gap Matters
In RAND's political-military games of the 1960s, observers repeatedly noted a phenomenon the staff called "doctrinal drift": players would describe their decision logic in pre-game interviews, then behave differently the moment the scenario generated real time pressure. The gap wasn't deception. Players genuinely believed their self-reports.
The same pattern showed up in the UK MoD's Foresight exercises in the early 2000s, where senior officials who described themselves as "comfortable with uncertainty" consistently demanded premature resolution of ambiguous intelligence before committing to any course of action. Their self-model said: tolerant of ambiguity. Their behavior said: desperate to close it.
If you never measure the gap, you can't design around it.
The Pre-Commitment Probe
One concrete technique: before the game starts, ask every player (in writing, individually) to answer three questions.
- What condition would cause you to recommend a major escalation in this scenario?
- What single piece of incoming intelligence would change your mind about the current course of action?
- If your team's plan is failing at the midpoint, what will you do?
Collect the answers. Give them to your umpire team. Do not show them to other players.
Then run the scenario. At the midpoint, have umpires privately compare each player's live behavior against their written answers. After the game, run a structured debrief where you read each player their own pre-commitments and ask them to account for the divergence.
The divergence is the data. Players who said they'd escalate on Condition X, then didn't when X arrived, have revealed something real about their actual decision calculus. That's worth forty minutes of debrief time.
Building the Mirror Into the Scenario
The pre-commitment probe is a bolt-on. A better approach builds the self-model test into scenario architecture from the start.
Design two parallel decision points that are structurally identical but emotionally different. For example: in Turn 3, players must decide whether to commit reserve forces when the situation is ambiguous but the stakes feel low. In Turn 7, they face the same logical decision, same probability distribution on outcomes, but the framing involves a public commitment their leadership has already made.
Players who behave differently in Turn 7 than Turn 3 have shown you that sunk-cost pressure overrides their stated logic. They will not know they've shown you this. The mirror is broken; they think they're looking at consistent reasoning.
graph TD
A[Pre-Game: Written Self-Model] --> B{Scenario Runs}
B --> C[Turn 3: Low-Stakes Analogue Decision]
B --> D[Turn 7: High-Stakes Structural Twin]
C --> E{Umpire: Behavior Match?}
D --> E
E --> F[Divergence Log]
F --> G[Post-Game: Mirror Debrief]
Keep the structural twinning invisible during play. If players notice they're being compared to themselves, they'll game the comparison.
The Debrief Is the Weapon
None of this matters without a debrief designed to make the data land. Show players the divergence without editorializing. Read back their pre-game answers. Show the decision log from the game. Ask them to explain the difference.
Resistance is normal and useful. Players will offer situational explanations: "The Turn 7 decision felt different because..." That explanation is itself data. Write it down. Run the same player through a different scenario six months later and check whether the pattern repeats.
Over multiple iterations, you build a behavioral profile that is far more reliable than any self-reported leadership assessment. The player who consistently abandons their stated de-escalation threshold when a superior is watching has told you something important about how they'll behave in a real crisis.
Designing for the broken mirror doesn't require psychological expertise. It requires pre-commitment capture, structurally twinned decision points, and a debrief that confronts divergence directly. The self-model gap is measurable. Start measuring it.
Get Uncertainty Game in your inbox
New posts delivered directly. No spam.
No spam. Unsubscribe anytime.