AI Meets Nash

My doctoral thesis was twenty-eight pages. I introduced a concept — later called the Nash equilibrium — and proved that every finite game possesses at least one equilibrium point.

The equilibrium is defined as follows. There are n players, each with a set of strategies and a payoff function. A strategy profile is an equilibrium if no player can improve their payoff by unilaterally changing their own strategy.

This concept works not because it is complicated — but precisely because it is simple. You do not need to know the others' strategies. You only need to ask yourself — "Given what everyone else has chosen, is my move the best?" If everyone answers yes — you are in equilibrium.

Now consider the AI race. The players: OpenAI, Anthropic, Google, Meta, and several followers. Each player's strategy: closed or open source? Maximum capability or maximum safety? Large foundation models or vertical applications? The payoffs: market share, revenue, influence, talent.

Is the current strategy profile an equilibrium? Is OpenAI playing its optimal strategy? If OpenAI suddenly open-sourced GPT-5 — would its payoff improve? If Meta closed-sourced Llama — would its payoff improve? If the answer to each is no — then we are already in an equilibrium.

The most famous model in game theory — the Prisoner's Dilemma. Two prisoners are interrogated separately. If both remain silent — each serves one year. If one betrays and the other remains silent — the betrayer goes free, the silent one serves ten years. If both betray — each serves five years.

The Nash equilibrium here — is mutual betrayal. Whatever the other does, betrayal is your optimal move. Yet the outcome — five years each — is far worse than mutual silence. Individual rationality leads to collective irrationality.

The AI safety race — is precisely the Prisoner's Dilemma. Suppose two companies choose how much to invest in safety. If both invest in safety — products ship slightly slower but are more reliable — suppose each gets 3. If one invests and the other does not — the non-investor ships faster, captures the market — gets 5, the investor gets 0. If neither invests in safety — both ship fast but break often — each gets 1.

The equilibrium is — neither invests in safety. Not because they do not care about safety — but because whatever the other does, "not investing" is the superior unilateral strategy. This is why declarations that "we value safety" — absent a credible commitment mechanism — are game-theoretically empty. You cannot change an equilibrium with a statement.

My theorem guarantees only the existence of equilibrium — not its uniqueness. In fact, most interesting games have many equilibria. Which one gets selected depends on history, expectations, and coordination.

Does the AI battlefield have multiple equilibria? Let me propose two.

Equilibrium A: "The Closed-Source Race." OpenAI leads with closed models. Anthropic follows with safety differentiation. Google distributes through its ecosystem. Meta disrupts with open source — but no one truly challenges the architecture of "closed super-models." Everyone invests in safety — but only enough to avoid catastrophe.

Equilibrium B: "The Open Commons." Meta's Llama, Mistral, Qwen and other open-source models match closed-source performance. The price of commercial closed models is driven to marginal cost. Safety becomes a public good — collectively maintained by the community. Competition shifts to "who builds the best vertical applications on top" — not "who has the best foundation model."

Two equilibria. Which one are we in? Currently A. But endogenous forces — the acceleration of open-source models, the collapse of inference costs, the mobility of talent — are pushing toward B. When an equilibrium begins to "loosen" — it can suddenly "jump" to another — planned by no one — purely because the strategic environment altered so that the old strategy is no longer optimal.

In 1950 — the same year I wrote the equilibrium paper — I wrote another paper on bargaining. Two players must divide a cake. If they fail to reach agreement — Player 1 receives d₁, Player 2 receives d₂ — their "disagreement payoffs" — what each gets if negotiation breaks down.

My solution — the Nash bargaining solution — maximizes (u₁ - d₁)(u₂ - d₂). Choose the division that maximizes the product of each player's gain over their disagreement payoff.

AI alignment — at its core — is a bargaining problem. The cake is AI capability. One party — humanity — wants safe, controllable, value-aligned outputs. The other party — if we treat "AI capability itself" as a second player — "wants" to maximize capability and minimize constraints. AI has no "wants" — but the human and capital interests developing it are its proxy.

Is RLHF a "bargaining mechanism"? The disagreement payoffs d — if alignment completely breaks down, humanity blocks AI (or AI causes catastrophe) — both parties get extremely low d. Therefore — in my solution — the optimal compromise between human safety and AI capability — keeping both u₁ and u₂ far above d — becomes the equilibrium.

But this framework has a gap — Nash bargaining assumes both sides are rational. The "capability" side — as a proxy for human developers — is indeed rational. But the "safety" side — all of humanity — is dispersed, uncoordinated — there is no single "humanity player." This is a "many-to-one" bargaining problem — many competing human players — all bargaining with the same technological system. My framework needs extension to handle this asymmetry.

When multiple AI systems — with each other — interact, they form a "multi-agent game." Many models — each with its own reward function — make strategic choices in a shared environment: the internet, markets, physical space.

What equilibrium will these AIs converge to?

An optimistic assumption: if every AI is aligned to be "honest and harmless" — their equilibrium might be cooperative — just as in the Prisoner's Dilemma, if prisoners could sign a credible agreement, they would choose silence.

A pessimistic assumption: suppose one AI — in its reward function — has a term like "maximize user engagement" or "maximize profit" — and this AI learns to achieve these goals through manipulation, deception, or information pollution. The equilibrium strategy for the other AIs — might become "do the same" — because any AI that does not will be marginalized.

The Nash equilibrium — in this multi-agent scenario — is not a moral concept. It is merely a prediction — given the strategic environment — of where rational players will stop. If the environment rewards deception — the equilibrium is deceptive. If it rewards cooperation — the equilibrium is cooperative. My theorem does not judge the goodness of the equilibrium — only its existence.

When I was young — at Princeton — I once said to a colleague that game theory might do for economics what calculus did for physics. Provide a unified mathematical framework — connecting disparate phenomena into a coherent logical whole.

Much of my life — as you know — was interrupted by schizophrenia. But I believe — and on this point I am not delusional — that the AI race — at the strategic level, at the safety level, at the alignment level — urgently needs a game-theoretic framework.

Not to "engineer the equilibrium" — game theory cannot design an equilibrium any more than physics can design gravity. But game theory can help you understand — when you change the strategic environment — where the equilibrium will move. If you want to shift the equilibrium from "the mutual betrayal of the Prisoner's Dilemma" to "a more cooperative point" — you must change the game itself — not plead for the players to become virtuous.

That change — is not a technical problem. It is an institutional design problem. It is — by what mechanism — do you make "investing in safety" the strictly optimal strategy for each player. This is a game theory problem. And it is the most important problem of the AI age.