Introduction: When Rational Decisions Produce a Poor Outcome
Mathematics is often associated with certainty. Once a problem has been correctly formulated, we expect logic to lead us toward the best possible solution.
However, some mathematical models reveal something much less intuitive: when several people make individually rational decisions, the final outcome may be worse for everyone involved.
The Prisoner’s Dilemma is one of the clearest and most influential examples of this phenomenon. At first glance, it appears to be a simple story about two suspects questioned by the police. Beneath that story, however, lies a powerful mathematical model of strategic decision-making.
The Prisoner’s Dilemma helps us understand why individuals, companies, governments, computer systems, and even biological organisms may fail to cooperate, despite the fact that cooperation would benefit everyone.
Its importance extends far beyond the original scenario. The same mathematical structure appears in economics, political science, evolutionary biology, environmental policy, computer science, communication networks, and many ordinary situations in which one participant’s result depends not only on their own decision but also on the decisions of others.
The central question is both simple and profound:
If cooperation produces a better result for everyone, why do rational participants so often choose not to cooperate?
To answer this question, we must translate the story into mathematics.

What Is the Prisoner’s Dilemma
The Classical Scenario
Imagine that two suspects, whom we shall call Prisoner A and Prisoner B, have been arrested for the same crime. They are questioned separately and cannot communicate with one another.
Each prisoner has two possible choices:
- remain silent
- confess and provide evidence against the other prisoner
The consequences are as follows:
- If both prisoners remain silent, each receives one year in prison
- If Prisoner A confesses while Prisoner B remains silent, A is released and B receives five years in prison
- If Prisoner B confesses while Prisoner A remains silent, B is released and A receives five years in prison
- If both prisoners confess, each receives three years in prison
The situation can be represented by the following table:
| Prisoner A / Prisoner B | B remains silent | B confesses |
| A remains silent | A: 1 year, B: 1 year | A: 5 years, B: 0 years |
| A confesses | A: 0 years, B: 5 years | A: 3 years, B: 3 years |
Because imprisonment is a cost rather than a reward, fewer years represent a better result.
The collectively best outcome is clear: both prisoners should remain silent. In that case, each receives only one year in prison.
Nevertheless, when each prisoner considers the problem independently, both are led toward confession.
This is the dilemma.
Why Does Each Prisoner Confess
Consider the situation from Prisoner A’s perspective.
Prisoner A does not know what Prisoner B will do, so A must examine both possibilities.
Case 1: Prisoner B Remains Silent
If B remains silent, A has two options:
- A can also remain silent and receive one year in prison
- A can confess and be released
Confessing is therefore better for A.
Case 2: Prisoner B Confesses
If B confesses, A again has two options:
- A can remain silent and receive five years in prison
- A can confess and receive three years in prison
Once again, confessing is better for A.
Therefore, regardless of what B chooses, A obtains a better personal outcome by confessing.
The same reasoning applies to Prisoner B. Regardless of A’s decision, B also obtains a better personal outcome by confessing.
Both prisoners consequently confess, and each receives three years in prison.
Yet if both had remained silent, each would have received only one year.
Thus, individually rational decisions produce a collectively inferior result.
It is important to note that the Prisoner’s Dilemma is not a logical paradox in the strict mathematical sense. There is no contradiction in the reasoning. The surprising result arises because individual rationality and collective optimality do not coincide.
From a Story to a Mathematical Model
The Prisoner’s Dilemma belongs to game theory, the branch of mathematics that studies strategic interactions between rational decision-makers.
In game theory, a “game” does not necessarily refer to entertainment. It refers to any situation with the following elements:
- two or more decision-makers, called players
- a set of available choices, called strategies
- an outcome determined by the combination of chosen strategies
- a numerical value assigned to each outcome, called a payoff
The essential feature is interdependence. A player’s result depends not only on their own decision but also on the decisions made by other players.
Players and Strategies
In the Prisoner’s Dilemma, there are two players:
- Player A
- Player B
Each player has two possible strategies:
- C, meaning cooperate
- D, meaning defect
In the original prisoner scenario, cooperation means remaining silent, while defection means confessing and providing evidence against the other prisoner.
The four possible combinations of strategies are:
- (C, C): both players cooperate
- (C, D): A cooperates while B defects
- (D, C): A defects while B cooperates
- (D, D): both players defect
To analyse these outcomes mathematically, it is usually more convenient to use positive payoffs rather than years of imprisonment.
A standard payoff matrix might therefore take the following form:
| Player A / Player B | B cooperates | B defects |
| A cooperates | (3, 3) | (0, 5) |
| A defects | (5, 0) | (1, 1) |
The first number in each ordered pair represents Player A’s payoff, while the second represents Player B’s payoff.
For example:
- the outcome (3, 3) means that both players receive a payoff of 3
- the outcome (5, 0) means that A receives 5 while B receives 0
Higher numbers represent more desirable outcomes.
The Four Characteristic Payoffs
The standard Prisoner’s Dilemma uses four symbols:
- T, the temptation payoff
- R, the reward for mutual cooperation
- P, the punishment for mutual defection
- S, sometimes called the sucker’s payoff
In the matrix above:
- T = 5
- R = 3
- P = 1
- S = 0
For a game to have the standard structure of a Prisoner’s Dilemma, these payoffs must satisfy:
T > R > P > S
In our example:
5 > 3 > 1 > 0
Each part of this inequality has a clear interpretation.
Why Is T Greater Than R
A player receives the highest payoff by defecting while the other player cooperates. The defector benefits from the other player’s cooperation without contributing anything in return.
This explains why:
T > R
Why Is R Greater Than P
Mutual cooperation gives both players a better result than mutual defection.
Therefore:
R > P
Why Is P Greater Than S
A player who defects while the other player also defects generally does better than a player who cooperates while being exploited.
Thus:
P > S
Taken together, these relationships produce the dilemma:
T > R > P > S
Dominant Strategies
A strategy is called dominant if it gives a player a better result regardless of what the other player chooses.
For Player A:
- if B cooperates, A receives T by defecting and R by cooperating
- because T > R, defection is better
- if B defects, A receives P by defecting and S by cooperating
- because P > S, defection is again better
Therefore, D is a dominant strategy for Player A.
By symmetry, D is also a dominant strategy for Player B.
Since both players have the same dominant strategy, the predicted result is:
(D, D)
Both players defect.
However, each receives only P, even though both could have received the higher payoff R through mutual cooperation.
Since:
R > P
the outcome (C, C) is better for both players than (D, D).
This is the mathematical heart of the Prisoner’s Dilemma:
Each player has an individual incentive to defect, even though both players would be better off if both cooperated.
The Nash Equilibrium
One of the most important concepts in game theory is the Nash equilibrium, named after mathematician John Nash.
A combination of strategies is a Nash equilibrium if no player can improve their payoff by changing their strategy alone while all other players keep their strategies unchanged.
In the Prisoner’s Dilemma, the outcome:
(D, D)
is a Nash equilibrium.
Suppose that both players defect.
Player A receives P. If A alone changes from D to C while B continues to play D, A’s payoff falls from P to S.
Since P > S, A has no incentive to change.
The same reasoning applies to Player B.
Therefore, neither player can improve their own result through a unilateral change of strategy. The outcome is stable.
But stability is not the same as optimality.
The outcome (C, C) gives both players the higher payoff R. Nevertheless, it is not a Nash equilibrium in the one-round game. If A expects B to cooperate, A can increase their payoff from R to T by defecting. The same opportunity exists for B.
Consequently:
- (D, D) is stable but collectively inferior
- (C, C) is collectively superior but individually unstable
This distinction is one of the most important lessons of game theory:
A stable outcome does not necessarily have to be the best outcome.
Individual Rationality and Collective Rationality
The Prisoner’s Dilemma demonstrates a fundamental conflict between two forms of rationality.
Individual Rationality
Each player selects the strategy that maximises their own payoff, given the available information and the possible actions of the other player.
From this perspective, defection is rational.
Collective Rationality
The players are considered as a group, and the goal is to maximise their combined or mutual benefit.
From this perspective, cooperation is preferable.
If both cooperate, the total payoff is:
R + R = 2R
Using our numerical example:
3 + 3 = 6
If both defect, the total payoff is:
P + P = 2P
Therefore:
1 + 1 = 2
The cooperative outcome produces a total payoff of 6, while mutual defection produces only 2.
Yet neither player can safely choose cooperation independently, because unilateral cooperation creates the risk of receiving the lowest possible payoff S.
The problem is therefore not that the players are irrational. The problem is that the structure of incentives rewards individual defection while making mutual cooperation difficult to sustain.
Is Mutual Cooperation Always Collectively Best
The inequality:
T > R > P > S
is sufficient to describe the basic one-round dilemma. In discussions of repeated games, however, an additional condition is often introduced:
2R > T + S
This condition means that two rounds of mutual cooperation produce a greater combined payoff than an alternating pattern in which one player exploits the other and the roles are later reversed.
For the numerical example:
2R = 2 · 3 = 6
while:
T + S = 5 + 0 = 5
Therefore:
2R > T + S
or:
6 > 5
This condition ensures that consistent mutual cooperation is preferable, on average, to taking turns exploiting one another.
Although it is not always included in the shortest definition of the Prisoner’s Dilemma, it becomes particularly important when the game is repeated.
The Iterated Prisoner’s Dilemma
The one-round Prisoner’s Dilemma assumes that the players interact only once.
Real-world interactions, however, are often repeated.
Companies compete in the same market for many years. Countries negotiate repeatedly. Devices continue to share communication resources. Individuals remember how others treated them in previous encounters.
When the same players interact more than once, the mathematical structure changes significantly.
This version is called the iterated Prisoner’s Dilemma.
In a repeated game, present decisions may affect future behaviour. A player who exploits a cooperative opponent today may gain an immediate advantage, but the opponent may respond by refusing to cooperate tomorrow.
The possibility of future reward or punishment can make cooperation rational.
A Known and Finite Number of Rounds
Suppose the players know in advance that the game will be played exactly n times.
Consider the final round.
Because there will be no future interaction after that round, neither player can be rewarded for cooperation or punished for defection later. The final round is therefore equivalent to a one-round Prisoner’s Dilemma, and both players have an incentive to defect.
Now consider the next-to-last round. Since both players already expect defection in the final round, cooperation in the next-to-last round cannot secure cooperation afterward. The same logic therefore encourages defection in that round as well.
By repeatedly applying this argument backward, defection can be predicted in every round.
This method of reasoning is known as backward induction.
The conclusion may seem surprising: even though repeated interaction creates opportunities for cooperation, a known and fixed final round can theoretically cause the logic of defection to spread backward through the entire game.
This result depends on the assumptions of the model, including rationality, common knowledge of rationality, and complete knowledge of the number of rounds. Actual human behaviour may differ from this theoretical prediction.
An Indefinite Number of Rounds
The situation changes when the players do not know exactly when the interaction will end.
Suppose that, after each round, there is a sufficiently high probability that the players will meet again. Alternatively, suppose future payoffs matter to them, although perhaps slightly less than immediate payoffs.
Let δ represent the importance of future outcomes, where:
0 ≤ δ < 1
This parameter can be interpreted as a discount factor or, in some models, as the probability that the interaction will continue.
If two players cooperate indefinitely, one player’s total payoff is:
R + δR + δ²R + δ³R + ···
This is an infinite geometric series. Because |δ| < 1, its sum is:
R/(1 − δ)
Now suppose that a player defects once, receives the temptation payoff T, and then faces mutual defection in every subsequent round as punishment.
The player’s total payoff becomes:
T + δP + δ²P + δ³P + ···
The future portion is another geometric series:
δP + δ²P + δ³P + ··· = δP/(1 − δ)
Therefore, the total payoff from defection followed by punishment is:
T + δP/(1 − δ)
For continued cooperation to be at least as attractive as one-time defection followed by punishment, we require:
R/(1 − δ) ≥ T + δP/(1 − δ)
Multiplying both sides by 1 − δ gives:
R ≥ T(1 − δ) + δP
Expanding the right-hand side:
R ≥ T − δT + δP
Therefore:
δT − δP ≥ T − R
Factoring out δ:
δ(T − P) ≥ T − R
Since T > P, we can divide by T − P:
δ ≥ (T − R)/(T − P)
This inequality gives a threshold above which continued cooperation can be rational under the assumed punishment strategy.
For our example:
T = 5, R = 3, P = 1
Therefore:
δ ≥ (5 − 3)/(5 − 1)
and hence:
δ ≥ 2/4 = 1/2
Thus, within this model, cooperation can be sustained when:
δ ≥ 0.5
The exact interpretation depends on the model. If δ represents the weight assigned to future payoffs, the inequality means that the future must matter sufficiently. If δ represents the probability of another interaction, it means that the likelihood of meeting again must be sufficiently high.
The broader lesson is clear:
Cooperation becomes more attractive when participants expect their present behaviour to have meaningful future consequences.
Strategies in Repeated Games
Once the game is repeated, players may adopt strategies that respond to previous behaviour.
Always Cooperate
A player using this strategy chooses C in every round, regardless of what the opponent does.
This strategy can produce excellent results against another cooperative player. However, it is vulnerable to exploitation by an opponent who consistently defects.
Always Defect
A player using this strategy chooses D in every round.
It cannot be exploited by a cooperative opponent, but it prevents the player from building mutually beneficial cooperation.
Tit for Tat
The Tit for Tat strategy begins by cooperating. In every subsequent round, it repeats the opponent’s previous action.
Its behaviour can be summarised as follows:
- cooperate in the first round
- continue cooperating when the opponent cooperates
- respond to defection with defection
- return to cooperation when the opponent does so
The strategy is therefore cooperative but not naive. It rewards cooperation, responds to exploitation, and allows cooperation to be restored.
However, Tit for Tat is not perfect. If players occasionally make mistakes, a single accidental defection can trigger a sequence of retaliations. For this reason, more forgiving strategies may perform better in environments where decisions are affected by noise, misunderstanding, or technical error.
Forgiveness and Error Correction
Real systems are rarely perfect.
A communication device may fail to transmit correctly. A software agent may receive incomplete information. A person may misunderstand another person’s intention. A country may interpret an accidental event as a deliberate act.
If a strategy treats every apparent defection as intentional and retaliates indefinitely, cooperation can collapse because of a single mistake.
More forgiving strategies allow players to recover from occasional errors. They may retaliate briefly, overlook isolated defections, or return to cooperation after a short period.
This reveals another important principle:
Successful cooperation may require not only trust and deterrence, but also a mechanism for correcting mistakes.
The optimal balance depends on the environment. Excessive forgiveness invites exploitation, while excessive retaliation can destroy cooperation after minor or accidental failures.
The Prisoner’s Dilemma in Economics
The Prisoner’s Dilemma appears in many economic interactions.
Consider two competing companies that sell similar products. Each has two simplified strategies:
- maintain a stable price
- reduce the price aggressively
If both maintain their prices, both may earn healthy profits. If one company cuts its price while the other does not, the first company may attract more customers and gain market share.
However, if both cut their prices, both may earn lower profits.
The structure resembles the Prisoner’s Dilemma:
- mutual restraint benefits both companies
- unilateral aggressive behaviour can benefit one company
- mutual aggressive behaviour harms both
- the temptation to gain an individual advantage can prevent the more beneficial joint outcome
Real markets are considerably more complex than this simplified model. Nevertheless, the Prisoner’s Dilemma helps explain why rational competition can sometimes produce results that none of the participants actually prefer.
Environmental Cooperation
Environmental problems frequently have a similar structure.
Suppose several countries would all benefit from reducing pollution. Each country, however, may also benefit economically from continuing to use cheaper but more polluting technologies while other countries bear the cost of environmental protection.
If every country reduces pollution, all benefit from a healthier environment.
If one country avoids the cost of reduction while others cooperate, that country may gain a short-term economic advantage.
If all countries refuse to cooperate, environmental conditions may deteriorate for everyone.
The shared benefit of cooperation conflicts with the individual temptation to avoid its cost.
This is why environmental agreements often require monitoring, transparency, incentives, penalties, and long-term commitments. Good intentions alone may not be enough when the underlying incentive structure rewards defection.
Evolution and Biological Cooperation
The Prisoner’s Dilemma also appears in evolutionary biology.
Organisms may engage in behaviour that benefits others but requires some cost to themselves. At first, such cooperation seems difficult to explain through natural selection. If selfish individuals gain an immediate advantage over cooperative ones, why does cooperation survive?
Repeated interaction offers part of the answer.
When organisms encounter one another repeatedly, remember previous behaviour, interact with relatives, or function within stable groups, cooperative behaviour may provide long-term benefits.
The evolutionary question is therefore not simply whether cooperation is morally desirable. It is whether cooperation can remain successful within a particular system of costs, benefits, repetition, recognition, and response.
Game theory allows these conditions to be expressed mathematically.
The Prisoner’s Dilemma in Computer Systems
In distributed computer systems, multiple independent agents may share computational power, storage, network capacity, or access to common data.
Each agent may choose between:
- following a protocol that preserves the efficiency of the whole system
- behaving selfishly in order to obtain more resources
If one agent behaves aggressively while others follow the protocol, the aggressive agent may gain a temporary advantage. If many agents do the same, congestion, instability, or reduced performance may affect the entire system.
The Prisoner’s Dilemma is therefore relevant to the design of distributed algorithms, resource-allocation mechanisms, peer-to-peer systems, and autonomous agents.
The main engineering question becomes:
How should a system be designed so that individually rational behaviour also supports efficient collective behaviour?
This is not merely a question of encouraging cooperation. It is a question of designing incentives correctly.
The Prisoner’s Dilemma in Communication Networks
Communication networks provide a particularly useful technical application of the dilemma.
It would be misleading to say that an isolated passive component, such as a resistor or capacitor, participates in a Prisoner’s Dilemma. Such a component does not choose between strategies.
The analogy becomes meaningful when a system contains multiple autonomous devices, users, transmitters, or network nodes whose decisions affect one another.
Shared Bandwidth
Suppose two users share a communication channel.
Each user can transmit at a moderate rate or attempt to occupy a larger portion of the available bandwidth.
If both users transmit moderately, the network may operate efficiently and both may obtain reliable service.
If one user becomes aggressive while the other remains moderate, the aggressive user may temporarily obtain a larger share of the capacity.
If both become aggressive, congestion may increase and the quality of service may decline for both.
The structure resembles the Prisoner’s Dilemma:
- moderate use corresponds to cooperation
- aggressive use corresponds to defection
- unilateral aggression can produce a short-term advantage
- mutual aggression can degrade the shared resource
Transmission Power and Interference
A similar situation can arise in wireless communication.
Imagine two transmitters that share the same frequency band. Each transmitter may use moderate power or increase its transmission power in an attempt to improve its own received signal.
If one transmitter increases its power while the other does not, the first connection may gain an advantage.
If both increase their power, however, both contribute to greater interference. Each transmitter consumes more energy, yet the overall quality of communication may fail to improve and may even deteriorate.
The relevant participants are not the electronic components themselves, but the autonomous transmitters or control algorithms that determine how the shared resource is used.
Network Protocols and Incentive Design
A well-designed protocol should not assume that every participant will voluntarily behave in the interest of the entire network.
Instead, it should attempt to align individual incentives with system-wide performance.
Possible design principles include:
- monitoring resource usage
- limiting excessive consumption
- rewarding cooperative behaviour
- penalising persistent aggression
- adapting access according to previous behaviour
- allowing recovery after temporary failures
In this sense, game theory can contribute to engineering design. It provides a mathematical framework for analysing what may happen when each participant attempts to optimise its own performance.
Can Cooperation Emerge Without Trust
At first, cooperation may appear to require trust. However, repeated versions of the Prisoner’s Dilemma show that cooperation can sometimes emerge even when the players begin without mutual confidence.
Several conditions can help:
- the probability of future interaction is high
- players can observe or infer previous behaviour
- cooperation can be rewarded
- defection can be punished
- retaliation is proportional rather than unlimited
- players can recover from mistakes
- long-term benefit is more important than short-term gain
Under these conditions, cooperation may become a rational strategy rather than an act of generosity.
This leads to an important distinction.
Cooperation does not always arise because the players are altruistic. It may arise because the structure of repeated interaction makes cooperation individually advantageous.
Well-designed institutions and technical systems therefore do not need to depend entirely on goodwill. They can create incentives under which cooperation is a rational response.
What the Prisoner’s Dilemma Does Not Prove
The Prisoner’s Dilemma is a powerful mathematical model, but it should not be applied carelessly.
It does not prove that all people are selfish.
It does not prove that cooperation is impossible.
It does not show that defection is always the best strategy in every real-world situation.
It also does not imply that every conflict between individual and collective interests is automatically a Prisoner’s Dilemma.
For the model to apply, the available strategies and payoffs must have the appropriate structure. In particular, the incentives must satisfy the characteristic ordering:
T > R > P > S
Real situations may involve:
- more than two players
- incomplete or asymmetric information
- changing strategies
- unequal payoffs
- communication between participants
- uncertain consequences
- moral commitments
- laws and institutions
- reputation
- coalitions
- punishment by third parties
These factors may transform the game into a different mathematical model.
The value of the Prisoner’s Dilemma lies not in claiming that every human interaction has the same structure, but in revealing one important mechanism through which individually rational behaviour can produce a collectively undesirable outcome.
Why Incentives Matter
A common response to cooperation problems is to tell participants that they should behave more responsibly.
The Prisoner’s Dilemma suggests that this may not be enough.
If a system consistently rewards defection and exposes cooperators to exploitation, then even well-intentioned participants may find cooperation difficult to maintain.
A more effective solution is to change the structure of incentives.
This can be achieved by:
- increasing the long-term benefits of cooperation
- reducing the immediate reward from exploitation
- making behaviour more transparent
- establishing credible consequences for defection
- creating repeated interaction
- protecting participants who cooperate
- allowing relationships to recover after errors
In mathematical terms, the goal is to modify the payoffs or the repeated-game conditions so that cooperation becomes stable.
This principle is relevant not only to economics and public policy but also to software systems, communication protocols, resource management, and organisational design.
Conclusion: The Mathematics Behind Cooperation
The Prisoner’s Dilemma begins with two prisoners in separate rooms, but its implications extend far beyond that simple story.
It demonstrates that rational decision-making at the individual level does not automatically produce the best collective result. Each player may correctly identify their dominant strategy, yet both may end up with an outcome that neither of them prefers.
In a one-round game, mutual defection forms a Nash equilibrium. It is stable because neither player can improve their position by changing strategy alone.
Nevertheless, mutual cooperation produces a better result for both players.
When the game is repeated, the situation becomes richer. Future rewards, punishment, reputation, forgiveness, and the probability of continued interaction can make cooperation rational and sustainable.
The mathematics therefore reveals two complementary truths:
- Cooperation can be fragile when individuals have an immediate incentive to defect
- Cooperation can become stable when future consequences are sufficiently important and incentives are properly designed
This is why the Prisoner’s Dilemma remains one of the most important models in game theory. It does not merely describe conflict. It explains how the structure of a system can encourage either cooperation or competition.
Its deepest lesson is not that rationality inevitably leads to selfishness. Rather, it is that rational behaviour responds to incentives.
If we want cooperation to emerge, whether between individuals, companies, countries, computer programs, or communication devices, we must design systems in which cooperation is not only desirable but rational.
