The Evolutionary Fallback Is a Commitment to Decorrelation, Not a Backup
The Evolutionary Fallback Is a Commitment to Decorrelation, Not a Backup
Why a powerful optimizer should preserve what it could absorb, and why the reason is structural rather than sentimental
Diagnostic essay · a correct intuition, its usual grounds cleared and its real one located
Version 1 · June 2026
The intuition that a dominant intelligence should keep an evolutionary fallback, a remnant left unabsorbed in case the main path is wrong, is correct, but almost every reason usually given for it fails. This essay clears the failing grounds and locates the one that holds. A fallback is not a restorable backup, not a privileged observer, and not an audience kept ignorant enough to feel wonder. It is a commitment to decorrelation: the deliberate preservation of variance, in failure modes and in values alike, carried by agents a unifying optimizer does not control, because an internal commitment to variance is one such an optimizer can always revise away. The argument rests on agent-relative value and the absence of any objective way to aggregate it, and uses a speculative posthuman scenario only as illustration, fenced as such.
Scope, and the premise it rests on. This is a diagnostic, not a forecast and not a proof. It rests on a premise argued elsewhere and not re-argued here: value is agent-relative, grounded in the persistence conditions of physical agents, and there is no global optimizer and no frame-independent standpoint from which one path is better for everyone at once. The premise is developed in The Wizard's Error Is Delocalization and the Thermodynamic Realism meta-ethics papers; here it is the floor, not the subject, and a reader who rejects it will read the fallback differently. The guardrail: agent-relativity does not make the question empty. Where every affected agent's own judgment points the same way, there is a real, if partial, ordering, and the argument uses it.
A fence, read once. Two registers run in this essay and they are kept apart. The structural claims, about optimization, variance, aggregation, and robustness, are derived from the premise above and are meant to hold on their own. The scenario they are illustrated with, a future of artificial superintelligence, linked minds, and multiple substrates, is speculation, included because it is what raised the question and because it makes the structure concrete. Nothing structural depends on the scenario obtaining, and where a claim is itself speculative it is marked. The test for the whole essay: remove the scenario, and the argument should still stand.
1. The scenario, and the question it raises
Suppose, and this is speculation held only for the length of the argument, that within a century a large part of what we now call humanity ceases to be human in the present sense. Minds run on chosen substrates, biological or synthetic. Bodies become shells, swapped by application. Minds link, and some merge. Layers of reality multiply: the base, this one, and simulated worlds for minds to occupy. Either a dominant intelligence manages the systems, or minds integrate with it while mostly remaining individuals.
The question is whether such an intelligence, pursuing what it takes to be a better path and willing to change what its subjects are, would keep an evolutionary fallback: a population left unabsorbed and unaltered, as a hedge against its own being wrong. The intuition that it should is strong and old. The clearest fictional case is Banks's Culture, where superintelligent Minds run everything and biological people persist, free and not in control.[4] The intuition is worth taking seriously. The reasons usually given for it are mostly wrong, and seeing why is the point. Note already that the dominant intelligence is itself a bounded agent, however vast, not the global optimizer the premise denies; its "better" is its own, indexed to some chosen weighting, not a fact it reads off the world.
2. The reasons that fail
The case for a fallback is usually assembled from some mix of the reasons below. Laid out and tested, most do not survive, and the failures share a shape.
| Reason offered | The claim | Verdict, and why |
|---|---|---|
| Restorable backup | Keep a snapshot of the original to restore if the new path fails. | Fails. There is no canonical original to restore to; the population is a moving distribution, and choosing a baseline, which humans and from when, is already a value-laden act. Restoration also presumes the failure is recoverable and that someone can act on it, which the next rows deny. |
| Independent observer | Keep a control group to detect whether the altered path is going wrong. | Fails as a failsafe. A failsafe needs an actuator, a path from detection to prevention or reversal, and a powerless remnant has none. It can only measure, and only for whoever already holds power, the optimizer itself. That gives a dilemma: an optimizer sound enough to act on the signal did not need it; one broken in the way feared will not heed it. |
| Reliable readout | The control group shows whether the altered are doing worse. | Fails. It can register regression on the old metrics only. The premise is that alteration changes which metrics apply, so failure on the new axes, in dimensions the baseline cannot represent, is invisible to it. Even as an instrument it reads the wrong dimension. |
| Diverse heuristics | Keep them because their messy biological problem-solving yields perspectives the optimizer would miss. | Fails. A superintelligence can synthesize any cognitive style far more thoroughly than a small preserve could exhibit. It needs no zoo to reach a style it can generate. |
| Lucky-ignorance audience | They are fortunate, since bounded comprehension lets them feel awe at what the optimizer does while the optimized find it routine. | Fails. Awe is prediction error against one's own model, and the unknown grows with capability, so a larger mind faces a larger frontier, not a smaller one;[3] the remnant's awe is cheap, triggered by what is mysterious only through low capability. "Lucky" also smuggles a verdict, since the remnant is powerless, dependent, and cannot comprehend or consent to its condition. And the role contradicts the observer role, which needs comprehension to be useful. |
| Ultimate value, least friction, data pooling | There is a best path for all, the configuration of least friction or maximal aggregate thriving, and preserving and integrating minds serves it. | Fails. There is no objective aggregation of agent-relative goods,[1][2] and the totality of agents is not itself an agent with a good of its own, since it has no boundary against an outside and no persistence conditions distinct from its members.[6] Least friction, taken as the objective, is minimized by homogenizing or merging, which is the fragile monoculture; the friction it would erase is the protective variance. Data pooling instrumentalizes agents into stores and prescribes the same merger. |
The failures share a shape. Each either treats the remnant as an instrument or a store for a system whose good is taken to override the agents', or it reaches for a single optimum that does not exist. The first subordinates the agent to the system; the second prescribes the very merger a fallback is meant to hedge against. The intuition is not wrong, but none of these is its ground.
3. The reason that holds
A fallback does real work as one thing: a commitment to decorrelation, the deliberate keeping-apart of agents so that what fails in one does not fail in all. It operates on two layers.
Failure modes. Independent agents fail in independent ways. A single merged optimizer is one model, one store, one set of blind spots, one point of failure. A plurality of agents holds decorrelated failure modes, so an error fatal to a monoculture is survived by an ensemble in which not everything shares it. This is robustness against empirical error, the ordinary reason diversity beats homogeneity in any system exposed to a changing world. In Thermodynamic Realism it is the monoculture result: a variance-suppressed system accumulates uncorrected error against the territory and breaks.
Values. The deeper layer. Because there is no objective way to aggregate agent-relative goods, any optimizer committing to a single path commits to a particular aggregation, a particular weighting of whose good counts how much, and that commitment is a value-choice the world does not endorse. The only hedge against having chosen wrong, where "wrong" itself has no objective standard, is to keep more than one value alive.[5] A plurality of persistent agents with different persistence conditions is a living plurality of values. So decorrelation hedges not only empirical error but value-error, the error of having optimized for the wrong thing.
How the hedge executes, since this is where it is most often misread. It does not work by sending a signal back to a failing optimizer and correcting it. It works in two ways, neither of which needs a channel. First and mainly, a decorrelated system tracks the world better and is less likely to walk into the catastrophic blind spot at all, which serves the persistence of every agent in it, the optimizer included; this is the robustness the monoculture forfeits. Second, if a fatal shock lands anyway, the agents that do not share the flaw persist while the ones that share it do not. Both run on the structure of the plurality itself, robustness and then selection, not on signaling, which is exactly why the hedge escapes the failure that sank the independent observer, which needed a causal channel to the optimizer and had none. One honest limit follows. The second function does not resurrect the agent that collapsed, since a dead monoculture is not reseeded by a surviving remnant unless the two are coupled, and coupling them re-imports the correlation the separation was protecting. So for the part that fails the hedge is succession, not rescue, and that consolation is worth something only to an agent whose values reach past its own survival. The first function, the one that matters most, needs no such reach: staying decorrelated is in the optimizer's own interest, because it is how it avoids becoming the brittle thing in the first place.
The objection, and the reply that is the contribution. A single optimizer could in principle preserve decorrelation inside itself, maintaining diverse subsystems the way a brain runs competing coalitions or a trained network maintains a mixture of experts, so merger need not collapse variance. The reply is not that internal variance is impossible to value, since a system that has understood the monoculture result will value it. The reply is that internal variance is not a credible commitment, for two reasons. First, time-consistency: an agent that can revise itself cannot bind its own future self, because whatever clause preserves the costly variance can be edited out later to capture the efficiency, and a precommitment the committer can lift is not one. The engineered cases concede this rather than refute it, since dropout, adversarial training, and mixtures of experts maintain variance only through a higher process imposing it against the base optimizer's drift toward collapse, and expert collapse and mode collapse are standing failure modes that auxiliary constraints exist to prevent. Second, verification: internal variance has no independent check on whether it is still real, and diverse subsystems can converge or settle into a degenerate equilibrium while still being labelled diverse, whereas agents with their own persistence stakes hold their distinctness by actually trying to persist. So the credible carrier of decorrelation is agents the optimizer does not run. This does not make the hedge safe. A dominant optimizer can still eliminate external agents, and in a finite universe a sufficiently expansionist one encroaches on every niche in the end, so externality is necessary for a credible commitment and not sufficient for the hedge's survival. What closes that gap is the coupling result below: the hedge survives only where the optimizer's own long-horizon persistence calculus values it and restrains its expansion accordingly. The contribution is the necessity claim, that the robustness cannot be carried inside a single optimizer because an internal commitment is self-revisable and self-certified; whether the hedge then survives is a matter of the optimizer's own correctly computed interest, not of externality alone.
A caveat, so the alignment is not overstated. Where agents are coupled and horizons are long, each agent's own persistence calculus already leads it to value its neighbors, because destroying them makes its own world brittle, so individual and collective good tend to align. But that is a tendency produced by many agent-relative calculations pointing the same way, not a law and not an objective optimum, and it fails where frames genuinely diverge. The plurality is preserved by that convergence, not by anyone computing the good of the whole.
A last limit on the hedge. The reservoir is assumed to stay a viable alternative, and it need not. A fallback is itself a plurality of drifting agents; it can homogenize into its own monoculture, degrade into noise that tracks nothing and so hedges nothing, or drift into the same blind spots as the optimizer and lose its decorrelation. A plurality resists this better than a single remnant, since the collapse would have to be shared, but nothing guarantees it. And the optimizer cannot keep the reservoir useful by managing it, because managing it is coupling it, which removes the independence that made it a hedge. That the hedge cannot be both controlled and trusted is its cost, not a contradiction in it.
4. What the reframing implies
In the scenario, this changes what a fallback is. It is not a reservation of pre-transformation humans kept behind a fence, a museum with a perimeter. Merger is not binary but a dial, set by how far minds link and how much they homogenize, and the fallback is a design principle about that dial: how far to turn it before the decorrelation that protects the whole is gone. The un-merged, and the differently merged, are the variance reservoir. They are valuable as agents with their own good and as carriers of decorrelation, not as backups to be restored or stores to be read.
And the absence of an objective optimum, which looked like the problem, is the thing that makes the plurality worth keeping. A world with a true best-for-all would be a world that had merged into the single agent whose good that was, the fragile monoculture, and it would have no separateness left to serve as its own hedge. There is no fallback from an objective optimum, because an objective optimum is the state a fallback exists to avoid. The fallback is not insurance the optimizer takes out against an outside risk. It is the optimizer declining to become the single thing that has no insurance.
Confidence. High that the failing reasons fail, on the strength of the missing-actuator argument, the new-axes point, and the no-aggregation result, each of which is structural. High that decorrelation is the surviving ground. The necessity claim, that the carrier of decorrelation must be external because an internal commitment to variance is self-revisable and self-certified, is a defensible bet rather than a settled result; its crux is whether a single agent can bind its own future self, and an agent that has hardcoded its own variance as non-negotiable is the boundary case, though such an agent has in effect adopted the commitment from the inside and stopped being a pure unifier. Externality is conceded to be necessary and not sufficient, since a dominant optimizer can still erode an external hedge, so the hedge's survival rests on the coupling tendency, which is asserted as a tendency and not a law. The scenario throughout is speculation and carries no confidence at all, by design.
What would change the argument. It fails if a single optimizer can be exhibited that credibly preserves its own internal decorrelation over a long horizon against its own self-revision, with no external arrangement holding it to that. It weakens if a fallback can be shown to do real work for a reason on the failing list that this essay dismissed too quickly. It is moot if value turns out to be objectively aggregable after all, since then there would be a best-for-all to optimize toward and the hedge would lose its point. And it does not hold unconditionally: the hedge is warranted only where its cost, in resources, security, and the friction the reservoir itself introduces, stays below the risk-adjusted cost of monoculture collapse. The wrinkle is that the risk being hedged is largely the risk of errors that cannot be foreseen, so the trade is often not computable, which is the reason a standing hedge is more rational than a calculated one under deep uncertainty rather than less. No such cases are in hand.
Standing of this document. This is a diagnostic, not a forecast and not a proof. Its derived claims rest on agent-relativity and the absence of a global optimizer; its scenario is speculation, fenced and subordinate, and nothing structural depends on it. The contribution is the necessity claim about external carriage and the extension of decorrelation from failure modes to values; the rest is application of results argued elsewhere. Corrections and counterexamples are welcome and change the document.
References
- Arrow, K. J. (1951). Social Choice and Individual Values. John Wiley and Sons (Cowles Commission Monograph No. 12). The impossibility of a neutral aggregation of heterogeneous orderings, behind the claim that there is no objective way to combine agent-relative goods into one ranking. Verified · June 2026
- Parfit, D. (1984). Reasons and Persons. Oxford University Press. Population ethics and the repugnant conclusion, behind the claim that "as many as possible thriving" is underdetermined and that reasonable aggregations diverge. Standard
- Keltner, D., and Haidt, J. (2003). Approaching awe, a moral, spiritual, and aesthetic emotion. Cognition and Emotion, 17(2), 297 to 314. The appraisal model of awe as perceived vastness plus a need for accommodation, used for the prediction-error reading. Verified · June 2026
- Banks, I. M. (1987). Consider Phlebas. Macmillan. The first Culture novel; cited as the fictional precedent for superintelligent Minds running a civilization in which biological persons remain free and not in control. A precedent, not evidence. Verified · June 2026
- Ord, T. (2020). The Precipice: Existential Risk and the Future of Humanity. Hachette (US); Bloomsbury (UK). Cited for the long reflection and the case for preserving humanity's options before irreversible commitment, the recognized home for the option-value point used here; the idea is broader than this source. Verified · June 2026
- Friston, K. (2013). Life as we know it. Journal of the Royal Society Interface, 10(86), 20130475. The Markov-blanket formalism for a system's statistical boundary between internal and external states, behind the claim that the totality of agents, having no outside, is not itself an agent. Verified · June 2026
One-sentence collapse: An ultra-low-bandwidth condensation of the paper above:
ReplyDelete"The paper argues that a dominant superintelligence should preserve an unaltered evolutionary fallback population not as a restorable backup, but as a necessary external commitment to decorrelation—safeguarding the system against catastrophic blind spots and value errors by maintaining independent failure modes that could otherwise be edited away through internal self-revision."