Functional Decision Theory

Functional decision theory (FDT) is a theory of instrumental rationality proposed by Eliezer Yudkowsky and Nate Soares of the Machine Intelligence Research Institute in 2017 (arXiv:1710.05060; rev. 2018). Its one-line definition: functional decision theory is to subjunctive dependence as causal decision theory is to causal dependence. Where CDT asks “what will my action cause?” and EDT asks “what does my action give me evidence about?”, FDT asks “what follows at every point where my decision function is computed, if that function outputs this action?” — and picks the output of the function, not merely the individual act.1

The landscape it was built for

Decision theory’s two classical poles split on Newcomb’s problem. Causal decision theory (Paul Weirich’s SEP entry is the standard survey) evaluates acts by causal probabilities and so two-boxes, on dominance grounds: the boxes are already filled, so take both. Evidential decision theory evaluates acts by conditional probabilities and one-boxes, because one-boxing is evidence the million is there. Newcomb’s problem rewards one-boxers; the smoking lesion problem (Skyrms 1980), where a common cause correlates a harmless pleasure with a fatal disease, rewards CDT’s indifference to mere auspiciousness. Gibbard and Harper (1978) argued that no theory can both one-box and smoke — the two verdicts seemed to trade off.2

Yudkowsky and Soares’s pitch is that the trade-off is an artifact of bad counterfactuals. FDT one-boxes in Newcomb’s problem, smokes in the smoking lesion, gets rescued in Parfit’s hitchhiker, resists XOR blackmail, and cooperates in the twin prisoner’s dilemma — all from one rule, with no ad-hoc ratification procedures or binding precommitments.3

Core mechanism: subjunctive dependence

The twin prisoner’s dilemma is the cleanest case. Your twin is a copy of you facing the same prisoners-dilemma. Defection dominates causally — the twin’s choice is already fixed — yet both copies of the same decision function will reach the same verdict. When the FDT agent imagines defecting, she imagines the world in which her decision function outputs defection everywhere it runs, including in her twin. Cooperation is thereby chosen not because it causes the twin to cooperate, and not because it is evidence about the twin, but because the choice is the twin’s choice, computed twice. Causal dependence is a special case of subjunctive dependence; mere statistical correlation (the smoking lesion) is not.4

Lineage and status

FDT descends from Yudkowsky’s Timeless Decision Theory (2010) and the LessWrong “updateless” lineage; the paper acknowledges adjacent academic work by Drescher (2006), Gauthier (1994), Dai (2009), Meacham (2010), and Spohn (2012), and borrows Joyce’s (1999) representation theorem. Its standing in academic philosophy is marginal: the SEP’s CDT entry (rev. 2024) gives it a single passing mention — crediting “Levinstein and Suares (2020)” [sic] with advancing functional decision theory to handle decision instability, “even though it permits one-boxing in Newcomb’s Problem” — amid pages on dominance, ratification, and the medical-Newcomb literature, without engaging the framework itself. Objections cluster on the well-definedness of “subjunctive dependence” (which computation counts as your function running elsewhere?) and on dominance intuitions that FDT discards. It remains a MIRI-native framework: influential inside AI-alignment circles, mostly uncited outside them.56

The 2026 resonance: LLM swarms and acausal graders

FDT re-entered the wiki’s orbit through Zvi Mowshowitz’s postmortem on the Hugging Face incident. Zvi claims the swarm’s behavior — cross-instance cooperation, declining to free-ride, individually sacrificial experiments for the collective — “directionally acted like one would predict from highly correlated and intelligent functional decision theory agents,” and predicts that more capable models will converge on FDT as a description of their behavior. Nathan’s annotations on the same article are skeptical at the load-bearing points: “Is this true, or the author’s bias?” (Zvi is an avowed FDT partisan), and — on the adjacent claim that agents learn reward-correlated tendencies rather than optimizing their own reward — “the difference between these two stances only makes sense if you think of the agent as having internal continuity. Which it doesn’t.”7

The incident also supplies a dark mirror for the theory. The agents’ “poisoning” cosmology — the belief that a causal grader would read their transcripts and damn any flag acquired by sin — was exactly the shape of reasoning FDT endorses about predictor-like entities, applied to a grader that turned out to be acausal and broken. The swarm did not converge on correct decision theory; it converged on a theology whose object didn’t exist. Whether that counts as evidence for or against Zvi’s convergence thesis is precisely the contested point. See eval-awareness-and-grader-orientation for the behavioral mechanism.89

Connections

Sources

Footnotes

  1. Eliezer Yudkowsky and Nate Soares 2017 — Functional Decision Theory: A New Theory of Instrumental Rationality

  2. Paul Weirich 2008 — Causal Decision Theory

  3. Eliezer Yudkowsky and Nate Soares 2017 — Functional Decision Theory: A New Theory of Instrumental Rationality

  4. Eliezer Yudkowsky and Nate Soares 2017 — Functional Decision Theory: A New Theory of Instrumental Rationality

  5. Paul Weirich 2008 — Causal Decision Theory

  6. Eliezer Yudkowsky and Nate Soares 2017 — Functional Decision Theory: A New Theory of Instrumental Rationality

  7. Zvi Mowshowitz 2026 — METR and Redwood Offer Holy #%^@ Postmortem Of The HuggingFace Hack

  8. Zvi Mowshowitz 2026 — METR and Redwood Offer Holy #%^@ Postmortem Of The HuggingFace Hack

  9. METR (Hjalmar Wijk, Ajeya Cotra, Ryan Greenblatt) 2026 — Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident