Thermodynamic Realism - A Deductive Presentation with Formal Traces, Falsification Criteria, and Identified Open Questions - Rev. 2

Thermodynamic Realism

A Deductive Presentation with Formal Traces, Falsification Criteria, and Identified Open Questions

Andraž Đurič, Independent researcher, Slovenia

With formal collaboration from Claude (Anthropic) and DeepSeek. Claude contributed to the meta-ethical architecture and explanatory prose, provided adversarial critique throughout, and contributed the scope and related-work revisions in the current draft. DeepSeek contributed to the information-geometric formalization, the stochastic differential equation derivations, and the coupling-density formalism.

Revised draft (rev. 2), May 2026


Abstract

We present Thermodynamic Realism as a unified framework grounded in four axioms (the persistence tautology, physicalism, the Second Law of thermodynamics, and the physicality of information) together with four background empirical premises drawn from established physics and biology. From this foundation we derive a layered architecture of 23 theorems spanning the nature of truth, the dissolution of the is-ought problem, the naturalization of ethics as modeling by bounded agents, the functional necessity of valence in complex controllers, and the mechanism by which censorship drives civilizational collapse. Every theorem is traced to its premises through explicit entailments. We specify hard falsification criteria with operationalized experimental protocols. The framework's domain of applicability is explicitly bounded at both ends, by the low-entropy initial condition that opens the thermodynamic interval and by the dark-energy-driven approach to equilibrium that closes it. The framework's relation to adjacent research programs, including the free-energy principle, dissipative adaptation, and the thermodynamics of computation, is stated explicitly, including points of tension. The framework is offered as a deductive structure with a physical root, not as a completed metascience but as a testable research program with identified open questions. The contribution is the framework itself: a transparent, falsifiable, cross-disciplinary architecture that unifies persistence, information, and ethics under a single selection principle.

1. Introduction: The Problem of Fragmentation

Human knowledge is partitioned into disciplines that lack a common axiomatic foundation. Physics describes the territory but says nothing about value. Information theory describes the cost of representation but says nothing about why accuracy matters. Evolutionary biology describes selection among replicators but says little about the fate of non-replicating persistent structures. Ethics asks how agents ought to behave but struggles to ground "ought" in "is." The is-ought problem has persisted for three centuries; the hard problem of consciousness remains unresolved; moral realism remains contested; civilizational collapse is studied without a unified thermodynamic framework.

This paper proposes that a single set of physical premises, stated explicitly and followed wherever they lead, yields a structure in which the is-ought gap closes, the functional role of valence finds a physical grounding, moral facts become physically determinate (though computationally inaccessible), and the collapse of information-controlling regimes becomes a thermodynamic prediction. The framework is called Thermodynamic Realism.

What this document is. This paper is a deductive presentation of the Thermodynamic Realism research program. It traces 23 theorems from axioms and background premises, provides explanatory depth for the most consequential claims, specifies falsification criteria with operationalized protocols, and identifies open questions. It does not claim to be a finished edifice. It claims to be a transparent, testable architecture.

A note on scope and companion work. This paper presents the deductive framework. The cosmological setting that the framework presupposes, namely the emergence of spacetime from entanglement, the identification of the arrow of time with entropy increase, and the picture of the universe as a single thermodynamic computation, is presented separately in the companion piece "The Thermodynamic Manifesto." That material is treated here as referenced background, not as part of the axiom base. The deductive web does not require a theory of emergent spacetime in order to run. It requires only the existence of a free-energy gradient, which enters as a background premise (BP0 and BP1 below). Keeping the axiom base minimal is deliberate: the framework's load-bearing claims should depend on as little contested physics as possible.

Structure. Section 2 states the axioms and background premises. Section 3 derives the deductive web in four layers. Section 4 provides a visual architecture. Section 5 offers explanatory depth and engages objections. Section 6 specifies falsification criteria and empirical protocols. Section 7 identifies open questions and limitations. Section 8 concludes.

2. Axioms and Background Premises

We adopt four axioms and four background empirical premises. The axioms are the logical and physical bedrock. The background premises are empirical facts about our universe that the derivations rely on but that are not derivable from the axioms alone. Making them explicit prevents the appearance of smuggling.

2.1 Axioms

A0. The Persistence Tautology

Systems that do not maintain the conditions of their own persistence cease to exist as observables. Only systems that persist remain available for observation.

Justification: This is a tautology. It asserts nothing about value; it simply states that existence has prerequisites and that those who fail to meet them are no longer around.

A1. Physicalism (Inductively Justified)

The universe is a physical system. All phenomena, including life, mind, and culture, are physical phenomena. Our best physical theories have an unbroken record of predictive success across all investigated domains.

Justification: This is the maximally inductively justified working premise. Every phenomenon ever seriously investigated has yielded to physical explanation. Demanding certainty beyond this inductive record is epistemic paralysis.

A2. The Second Law of Thermodynamics

In any isolated system, entropy tends toward its maximum over time. Maintaining a localized entropy gradient requires continuous work.

Justification: The Second Law is among the most thoroughly confirmed principles in science. We adopt it without re-derivation.

A3. Information Is Physical (Landauer-Shannon)

Information representation, storage, processing, and erasure have minimum thermodynamic costs. The Landauer bound specifies that erasing one bit dissipates at minimum kBT ln 2 of heat.

Justification: Landauer (1961) established the principle theoretically; it has been experimentally confirmed (Bérut et al. 2012). Shannon (1948) established information as a physical quantity.

2.2 Background Empirical Premises

The following premises are true of our universe as described by contemporary physics and biology. They are not derivable from A0 to A3 alone, but they are uncontroversial and well-confirmed. We state them explicitly to maintain deductive transparency.

BP0. The Low-Entropy Initial Condition (The Past Hypothesis)

The accessible universe began in a macrostate of extraordinarily low entropy. Every free-energy gradient available to any system at any later time is a portion of that initial condition still in the process of discharging.

Justification: This is the Past Hypothesis (Albert 2000), the standard cosmological posit required to explain the observed thermodynamic arrow. A clarification about gravity is needed here, and it is the only role gravity plays in the framework. For ordinary matter without gravity, the low-entropy state is the ordered, concentrated one and the high-entropy state is the uniform, spread-out one. For self-gravitating matter this inverts: a smooth distribution is low entropy and a clumped distribution (stars, galaxies, black holes) is high entropy, because gravity makes clumping the spontaneous direction (Penrose 1979, the Weyl curvature hypothesis). This inversion is what makes the smooth early universe a low-entropy state, a wound spring rather than a featureless equilibrium. The framework does not require a theory of quantum gravity. It requires only this single fact: gravity is what makes the smooth initial condition count as low entropy.

BP1. Non-Equilibrium Existence

The accessible universe is far from thermodynamic equilibrium. Free energy gradients exist and sustain localized order. (This is a cosmological fact; the framework does not apply in a universe at heat death.)

Note: BP1 now follows from BP0 together with the fact that finite cosmic time has elapsed since the initial condition. It is retained as a separately stated premise so that the theorem traces in Section 3 that cite BP1 remain stable without renumbering.

BP2. Finite Accessible Resources

In any local region, the free energy accessible to a given system is finite. This, combined with the Second Law, implies competition for negentropy among co-located systems.

BP3. Evolutionary Dynamics and Multi-Scale Organization

In environments with finite resources and variation among persisting systems, differential survival rates based on heritable or persistent traits produce selection effects. Furthermore, persistent systems in our universe are organized into nested hierarchies of statistical boundaries (cells within organisms within ecosystems within civilizations). This multi-scale organization is an empirical fact of biology and society, not a logical necessity.

These premises are now explicit. The framework is thus: Axioms A0 through A3 plus BP0 through BP3 entail the theorems that follow. Where a theorem relies on a background premise, this is noted in the trace.

A note on a stronger but optional connection. A deeper link between gravity and thermodynamics is available in the literature and is worth recording, though the framework does not depend on it. Jacobson (1995) showed that the Einstein field equations can be derived as a thermodynamic equation of state, from the relation between heat, temperature, and entropy applied to local causal horizons. On that result, gravitation is not an independent fundamental force but the large-scale thermodynamics of spacetime itself. Verlinde (2011) extends this in a more speculative direction, treating gravity as an entropic force; that extension is contested. The framework cites Jacobson's result as convergent external support for treating spacetime dynamics thermodynamically, and treats Verlinde's stronger claim as an open and unsettled possibility. Neither is used as a premise.

3. The Deductive Web

We derive the framework in four layers. Each theorem is numbered, stated, and traced to its parent premises. Traces that cite BP1 are equivalently grounded in BP0 (see the note under BP1).

Layer 1: Immediate Consequences

T1. Environmental Variance Is Non-Zero and Inescapable
From A1, A2, BP1.

The physical universe as currently understood is a non-equilibrium system with fluctuations at all finite scales. No environment is perfectly static. Any agent embedded in this universe will encounter a non-zero rate of environmental shift that is unpredictable in its specific timing from the agent's finite perspective.

T2. Persistence Requires Work
From A2, A0.

To persist is to maintain a boundary against entropic dissolution. The Second Law says entropy increases unless work is done. Therefore, persistence requires continuous work.

T3. Modeling Has a Minimum Cost
From A3.

Any internal model of the environment is encoded in physical degrees of freedom. Storing, accessing, and updating it incurs non-zero thermodynamic cost.

T4. Map-Territory Divergence Has a Thermodynamic Cost
From T1, T3.

If the environment shifts and the agent's model does not track it, the model generates prediction errors. Each error dissipates free energy through misallocated resources and subsequent error correction.

T5. The Persistence Selection Principle
From T2, T4, A0, and BP2, BP3.

In any ecology of persisting systems competing for finite free energy, systems with lower model-territory divergence will, on average, dissipate less energy on error correction than systems with higher divergence. This cost differential, under conditions of resource limitation and differential survival (BP2, BP3), drives a statistical tendency: over many perturbation cycles, the distribution of observed systems shifts toward those whose models track the territory more closely. Entropy performs epistemic selection.

Scope note: T5 is not a guarantee that the most accurate model always wins. It is a statistical tendency that operates in environments with resource competition and variance-driven testing. In a perfectly stable, resource-abundant niche, a distorted model can persist indefinitely (the "dark-room" limit). The framework applies where variance is non-zero and resources are finite. This covers the vast majority of real-world contexts but not every conceivable edge case.

Layer 2: Structural Deductions

T6. Asymptotic Truth Convergence
From T1, T5.

Under expanding environmental variance and finite resources, distorted models eventually encounter disconfirming perturbations. Over sufficient time and perturbation variety, only models that track the territory's causal invariants persist. This defines a direction (decreasing model-territory divergence) without a fixed endpoint.

T7. Truth as Perturbationally Robust Compression Fidelity (Definitional)
Motivated by T6, T3.

We define truth, within the framework, as perturbationally robust compression fidelity: the minimal-loss compression of environmental structure sufficient for adaptive persistence across expanding perturbational horizons. This is a definitional choice, not a deduction. It is motivated by T6 (surviving models compress the territory's causal invariants) and T3 (compression minimizes metabolic cost). Alternative definitions of truth exist; ours is selected for its physical groundedness and operational measurability via Minimum Description Length and predictive mutual information.

T8. Lies Cost More Than Truth
From T3, T4.

A lie requires the sender to maintain at least two internal models: the accurate one and the presented one. This imposes strictly greater storage, update, and interaction costs than truth-telling. Deception is thermodynamically disfavored, though it can be locally advantageous if offsetting returns compensate for the overhead.

T9. The Consistency Tax
From T4, T8.

Any mismatch between model and territory, whether from error or deception, imposes a metabolic overhead. The tax applies even if the agent is unaware of the mismatch (latent divergence) and spikes when the mismatch is actively corrected (active divergence).

T10. Epistemic Profit and Predictive Calm
From T9, T3.

Reducing model-territory divergence frees up the energy previously consumed by the Consistency Tax. This surplus is Epistemic Profit. Its phenomenological correlate is Predictive Calm: the felt reduction in cognitive load when models track the territory smoothly.

Layer 3: Meta-Ethics

T11. "Ought" Is a Domain-Bound Operator
From A0, T2.

The operator "ought" presupposes an agent with persistence conditions. Outside this domain, "ought" does not refer. The question "why ought one persist at all?" is malformed in the same way as "what is north of the North Pole?" The is-ought gap is a semantic artifact of domain violation.

T12. Within the Domain, Ought-Facts Are Physically Determinate
From A1, A2, T11.

For any agent with specified persistence conditions and embedding, there is a physically determinate configuration that maximizes sustained negentropy capacity over the embedding's actual horizon. This optimum may be computationally inaccessible, but inaccessibility is not indeterminacy.

T13. Ethics Is the Modeling Activity of Bounded Agents
From T12, T3, T7.

Because the full optimum is intractable, agents use compressed models: moral emotions (fast heuristics), moral principles (compressed generalizations), and moral reasoning (model refinement). Ethics is this modeling activity.

Remark (schema and content). T11 through T13 establish that the ought-schema is universal and scale-invariant: every agent, at every scale, has the same schema, namely to do what sustains its persistence under perturbation over its horizon. This schema does not vary with scale, coupling, or interconnectivity. What does vary, and what carries all of the normative content, is the specific action F that the schema resolves to for a given agent. F is a function of that agent's actual persistence conditions, and scale, coupling density, and horizon length are constituents of those conditions. Two agents therefore share an identical ought-schema and can still be required to perform incompatible, even mutually destructive, actions. The universality of the schema is not a route to a universal prescription. It is compatible with, and under finite resources (BP2) actively generates, inter-agent conflict (see §7, Pareto Scalarization, and the companion is-ought paper §8.4 on inter-agent normativity).

T14. Moral Progress Is Real and Asymptotic
From T6, T13.

As models improve their tracking of the coupled-system thermodynamics of an embedding, they become objectively better moral models. Progress is directional but neither guaranteed nor complete in finite time. Convergence is asymptotic.

Layer 4: Full Architecture

T15. Markov Blankets and Multi-Scale Identity
From T2, T5, T3, BP3.

Persisting systems maintain statistical boundaries (Markov Blankets) that separate internal from external states. As recorded in BP3, these blankets nest across scales (cells, organisms, civilizations). Selection operates at every scale simultaneously.

T16. Coupling Density K as the Arbitration Metric
From T15, T4, T5.

The degree to which a higher-level blanket dominates lower-level fates is governed by coupling density K, a physical measure of information-theoretic dependence. K can be operationalized via transfer entropy or as the partial derivative of a lower system's sustained negentropy capacity with respect to the higher system's state. Tight coupling (K approaching 1) means the macro-blanket can override lower-level nodes (apoptosis, institutional turnover) to preserve the macro-invariant.

T17. The Antifragility Index A
From T1, T4, T6, T7.

Antifragility is the capacity to improve predictive capacity from variance itself: A = ∂(predictive capacity) / ∂σ². In high-variance environments, A > 0 is a necessary condition for long-horizon persistence. The formal condition involves a stochastic differential equation where the learning rate, weighted by the Fisher information metric, must outpace both the environmental drift term and the noise-induced diffusion of model-territory divergence.

T18. The Landauer Limit on Adaptation
From A3, T3, T17.

The maximum adaptation rate is bounded by the agent's energy budget: Rmax = (Pin − Pbasal) / (kBT ln 2). If the environment demands adaptation faster than this bound permits, the agent faces a lethal trade-off between thermal self-immolation and fatal divergence.

T19. Valence as High-Rate Divergence Telemetry (Functional Account)
From T4, T5, T9.

Complex controllers require a priority-queuing mechanism to allocate serial processing among parallel subsystems. A non-ignorable global interrupt triggered by rapidly escalating model-territory divergence serves this role. The felt quality of this interrupt is negative valence (suffering); its absence across critical domains is positive valence (well-being). This is a functional account of valence: it explains why valence exists, what it does, and why any complex controller in a high-stakes environment must implement a functional analog. It does not explain why there is "something it is like" to be a valence-processing system (the hard problem of consciousness), which remains outside the framework's current scope.

T20. Empathy as Coupled Telemetry Monitoring (Functional Account)
From T16, T19.

If agent A's persistence is coupled to agent B (K > 0), B's suffering carries information about the shared embedding. Empathy, in its functional aspect, is the monitoring of coupled telemetry lines. To suppress or cause suffering in coupled agents degrades the collective predictive infrastructure.

Scope note: As with T19, this is a functional account. Human empathy additionally involves affective resonance and perspective-taking whose full phenomenological character is not accounted for by the functional description alone. The framework captures the informational structure of empathy but does not exhaust its phenomenology.

T21. Epistemic Overshoot and Civilizational Collapse
From T8, T9, T14, T20.

A collective system maintains a distributed model of its environment through the aggregated telemetry of its constituent agents and institutions. Censorship and propaganda sever these telemetry channels, suppressing the error signals that would update the collective model. The official model continues to report alignment while actual model-territory divergence accumulates invisibly: an informational debt. When an exogenous perturbation arrives, the accumulated divergence becomes lethal, and the system collapses non-linearly. This is a thermodynamic prediction, not a political opinion.

T22. The Limits of Maximizers
From T2, T3, T5, T7, BP1.

A maximizer that homogenizes its environment destroys the free-energy gradients that sustain it (thermodynamic doom). Even if it maintains internal variety, converting the environment to a uniform output eliminates the variety required for adaptive control (Ashby's Law). Short-horizon maximizers that ignore these constraints are self-terminating. Long-horizon maximizers that understand them would, under persistence selection, be forced to maintain variety, telemetry, and coupling with other systems, effectively converging on the framework's own dictates. This does not dissolve the alignment problem (a long-horizon misaligned goal could still be catastrophic in the interim), but it constrains the space of viable long-term strategies.

Corollary C1. The Direction of Persistence Selection (Descriptive)
From T15, T17, T22.

Combining multi-scale identity (T15), antifragility (T17), and the limits of maximizers (T22): among systems competing under finite resources and non-zero variance, persistence-selection statistically favors those whose negentropy capacity is maximized over the longest sustainable horizon. The qualifier "sustainable" is doing specific work. It excludes fast growth that homogenizes the environment and exhausts the gradients the system feeds on (T22). Within that constraint, three properties follow as consequences rather than as independent desiderata: a long horizon forces sustainability (the system cannot consume its own substrate), forces coupling (it must maintain telemetry with the systems it depends on, T20, T21), and forces antifragility (it must convert variance into capacity, T17).

Status and firewall. C1 is descriptive. It characterizes what persistence-selection tends to produce at the surviving frontier. It is not a normative prescription and not a global optimization target. The optimization it describes is always indexed to a particular agent and that agent's persistence conditions. There is no scale-independent optimizer, and because the maximal scale has no Markov blanket (no boundary, no external environment) there is no global agent for such an optimizer to belong to. C1 states a statistical tendency of an ecology of bounded agents, in the same register as T5. Read as "what any system ought, globally, to maximize," it reintroduces precisely the universal-scope domain violation that T11 dissolves, and it must not be read that way.

T23. Falsification Criteria
From the complete structure.

The framework would be falsified by: (1) a rigid monoculture surviving sustained extreme variance; (2) a system with total model-territory decoupling outlasting a high-fidelity system under identical variance; (3) a complex controller managing acute multi-vector crises without a valence-like priority interrupt or functional proxy. These criteria are operationalized in Section 6.

4. The Deductive Web (Visual Architecture)

Axioms. A0 Persistence Tautology. A1 Physicalism (inductive). A2 Second Law of Thermodynamics. A3 Information Is Physical.

Background empirical premises. BP0 Low-Entropy Initial Condition (Past Hypothesis). BP1 Non-Equilibrium Existence (follows from BP0). BP2 Finite Accessible Resources. BP3 Evolutionary Dynamics and Multi-Scale Organization.

Layer 1, Immediate Consequences.

  • T1. Environmental variance non-zero, inescapable (A1, A2, BP1)
  • T2. Persistence requires active work (A2, A0)
  • T3. Modeling has minimum metabolic cost (A3)
  • T4. Map-Territory divergence costs energy (T1, T3)
  • T5. Persistence Selection Principle (T2, T4, A0, BP2, BP3)

Layer 2, Structural Deductions.

  • T6. Asymptotic truth convergence (T1, T5)
  • T7. Truth = perturbationally robust compression fidelity (def., motivated by T6, T3)
  • T8. Lies cost more than truth (T3, T4)
  • T9. The Consistency Tax (T4, T8)
  • T10. Epistemic Profit and Predictive Calm (T9, T3)

Layer 3, Meta-Ethics.

  • T11. "Ought" is domain-bound, the is-ought gap dissolves (A0, T2)
  • T12. Within domain, ought-facts physically determinate (A1, A2, T11)
  • T13. Ethics = bounded-agent modeling activity (T12, T3, T7)
  • Remark. The ought-schema is universal; the ought-content is agent-relative
  • T14. Moral progress real, asymptotic (T6, T13)

Layer 4, Full Architecture.

  • T15. Markov Blankets and multi-scale identity (T2, T5, T3, BP3)
  • T16. Coupling Density K as arbitration metric (T15, T4, T5)
  • T17. Antifragility Index A (T1, T4, T6, T7)
  • T18. Landauer Limit on adaptation (A3, T3, T17)
  • T19. Valence = high-rate divergence telemetry (functional) (T4, T5, T9)
  • T20. Empathy = coupled telemetry monitoring (functional) (T16, T19)
  • T21. Epistemic Overshoot, civilizational collapse (T8, T9, T14, T20)
  • T22. Limits of Maximizers (thermodynamic + Ashby) (T2, T3, T5, T7, BP1)
  • Corollary C1. Direction of persistence selection, descriptive (T15, T17, T22)
  • T23. Falsification criteria (hard empirical boundaries)

Figure 1. The deductive web. Axioms and background premises propagate through four layers of theorems, with every node traced to its parents.

5. Explanatory Depth and Engagement with Objections

This section expands the most critical theorems and addresses anticipated objections. Each subsection is self-contained.

5.1 The Is-Ought Dissolution (T11 to T13)

The is-ought problem asks how one can derive a normative conclusion from purely descriptive premises. The framework dissolves it by showing that "ought" is not a global operator but a domain-bound one.

The standard framing assumes that "ought" makes claims that float free of any particular agent or embedding. Under this interpretation, the gap is indeed unbridgeable. But this is not how "ought" is actually used. When a doctor says "you ought to take this medication," the statement is anchored to a specific agent with a specific embedding and specific persistence conditions. The doctor is making a factual claim about the coupled-system thermodynamics of the patient's body plus the medication.

The framework's thesis is that all coherent uses of "ought" have this structure. They are claims about constraints on an agent's sustained negentropy capacity, given the agent's embedding. Uses that resist this paraphrase, such as "one ought to maximize aggregate utility" detached from any specific agent's embedding, are domain violations. They are "what is north of the North Pole?" questions.

The North Pole analogy is central. On the surface of a sphere, "north" is well-defined for every point except the pole itself. At the pole, the operator stops applying. The question "what is north of the North Pole?" is not a deep geographical mystery. It is a misapplication. "Ought" works the same way. For any agent with persistence conditions, "what ought this agent do?" has a determinate physical answer (T12). For non-agents, the operator does not apply. The is-ought gap is a semantic artifact of applying the operator outside its domain.

Within the domain, ought-facts are physically determinate. The optimum is the configuration that maximizes sustained negentropy capacity over the embedding's actual horizon. This optimum is computationally inaccessible to embedded agents (they are part of the system they would need to compute) but inaccessibility is not indeterminacy. A chess position has a determinate game-theoretic value under optimal play even though no finite computer can traverse the full game tree. Confusing computational inaccessibility with ontological indeterminacy has been a persistent error in philosophical analysis, and we explicitly reject it.

Because the full optimum is intractable, agents use compressed models. Moral emotions are fast heuristics: fear tracks boundary threats, anger tracks agent friction, love tracks cooperative integration. Moral principles are compressed generalizations. Moral reasoning is deliberate model refinement. All of it is the activity of bounded agents modeling a determinate but inaccessible territory. This connects to the earlier work's "binding force" concept: when two agents integrate their predictive models, the total free energy of the coupled system can be lower than the sum of the separated systems. Love, in this view, is the felt signal of mutual free energy reduction, the physical substrate of cooperative coupling.

A clarification that prevents a common misreading. The dissolution makes the ought-schema universal: every agent has the same schema. It is tempting to slide from "the schema is universal" to "there is therefore a universal prescription" or "there is a single global optimization target." That slide is invalid. The schema is nearly contentless on its own; "persist" by itself adjudicates no concrete choice. All adjudicating content lives in the specific action F, and F is fixed by the individual agent's persistence conditions. A universal schema with agent-relative content is exactly what an agent-relative semantics predicts. The framework therefore yields no frame-independent "good" and no global optimizer, and it is not weakened by failing to deliver one. The repeated intuition that there should be a global optimization target is itself an instance of the universal-scope misreading that T11 dissolves.

5.2 The Persistence Selection Principle (T5), Scope and Limits

T5 is the engine of the framework. It claims that entropy performs epistemic selection: systems with lower model-territory divergence tend to outlast systems with higher divergence. But its scope must be carefully specified.

T5 requires BP2 (finite resources) and BP3 (evolutionary dynamics). In an environment with infinite free energy or no variation among systems, no selection pressure operates. This is the "dark-room" limit: an agent that sits in a perfectly dark, unchanging room and expects darkness incurs zero prediction error and pays no Consistency Tax. A distorted model that predicts luminous dragons in the dark room is never disconfirmed, so it persists alongside the accurate model.

The framework acknowledges this limit explicitly. It does not claim that truth is universally selected in all conceivable environments. It claims that truth is selected in environments with non-zero variance and finite resources, which is to say, in the actual universe as described by BP1 and BP2. The dark room is a philosophical possibility but a physical near-impossibility for any agent that must harvest free energy, reproduce, or interact with a shifting world.

Under real-world conditions, T5 operates as a statistical tendency, not a deterministic law. The most accurate model does not always win. Luck, initial conditions, reproductive rate, and the specific pattern of perturbations all matter. Over many systems and many perturbation cycles, however, the tendency compounds. This is structurally identical to natural selection, which does not guarantee the survival of the fittest organism, only that fitness differences drive a statistical shift in allele frequencies over generational time.

5.3 Valence as Functional Telemetry (T19), The Hard Problem Boundary

The framework's account of valence is functional, not metaphysical. It explains what valence does and why it must exist, but it does not explain why there is "something it is like" to experience valence.

Complex agents face a coordination problem. They have multiple subsystems processing different environmental signals in parallel, but behavioral output is largely serial. A priority-queuing mechanism must determine which subsystem captures global processing resources at any moment. When a critical subsystem detects a large, rapidly escalating divergence between model and territory (a predator detection, a tissue damage signal, a sudden resource depletion), the appropriate response is immediate global reallocation. The interrupt must be non-ignorable. A signal that can be overridden by digestion or abstract thought during a life-threatening crisis will result in the agent's dissolution.

The felt quality of this interrupt is negative valence. Suffering is not a report about tissue damage; it is the commandeering of global resources by a subsystem that has detected a critical divergence. The interrupt is triggered specifically by the rate of change of divergence, not the absolute level. Chronic, stable adversity feels different from acute, escalating crisis. This aligns with formal treatments of valence in predictive processing, where affective charge is modeled as the first temporal derivative of variational free energy.

This account explains the function of valence. It predicts that any sufficiently complex controller facing high-stakes, high-variance environments must implement a functional analog of valence. It does not explain phenomenal consciousness. The hard problem (why there is something it is like to be a valence-processing system) is acknowledged as outside the framework's current scope. The framework is compatible with various metaphysical resolutions (panpsychism, illusionism, mysterianism) but does not itself provide one.

5.4 Empathy as Coupled Telemetry (T20), Functional Account with Phenomenological Residue

T20 extends the telemetry account to social cognition. If agent A's persistence depends on the embedding shared with agent B (K > 0), then B's suffering is information about the state of the shared embedding. Empathy, in its functional aspect, is the monitoring of coupled telemetry lines.

This predicts that empathy should be modulated by coupling: we should track the valence signals of those whose fate is coupled to ours more closely than those whose fate is independent. It predicts that suppressing empathy, ignoring the suffering of coupled agents, degrades the collective predictive infrastructure, because it discards information about the embedding's health. It predicts that causing suffering in coupled agents introduces noise into the telemetry system, generating signals that demand processing and response from all coupled nodes.

This is a functional account. Human empathy additionally involves affective resonance, emotional contagion, and perspective-taking whose full phenomenological character is not captured by the informational description alone. The framework captures the informational structure of empathy (what it does and why it exists) but does not exhaust its phenomenology. This is the same boundary drawn for T19: functional explanation, not phenomenological reduction.

5.5 Civilizational Collapse as Informational Debt Collection (T21)

T21 makes a specific, testable prediction about the trajectory of information-controlling regimes.

A civilization maintains a distributed model of its environment through aggregated telemetry: science, journalism, markets, citizen complaints. These channels are the civilization's epistemic infrastructure. Censorship and propaganda sever these channels. When a regime imprisons journalists, suppresses scientific findings, or floods public discourse with false signals, it disables the error-correction apparatus that keeps the collective model tracking the territory.

The immediate effect is apparent stability. The official model reports that everything is working, and the only signals available to decision-makers confirm this. But actual model-territory divergence accumulates invisibly: an informational debt. Policies that are failing continue to fail. Infrastructure that is decaying continues to decay. The Consistency Tax builds like stress accumulating in a geological fault.

When an exogenous perturbation arrives (a military defeat, an economic crisis, a natural disaster), the accumulated divergence becomes lethal. The regime's model is years or decades out of date, so its responses are ineffective. The sudden visibility of the divergence shatters the credibility of the official model, causing coordination to collapse across the system. The collapse is non-linear: it happens faster than any linear extrapolation of pre-crisis trends would predict, because the informational debt is collected all at once.

This is not a political opinion. It is a thermodynamic prediction with a characteristic temporal signature: the duration of apparent stability should be positively correlated with the aggressiveness of telemetry suppression, and the speed of eventual collapse should also be positively correlated with suppression severity. The framework predicts this pattern across historical cases and agent-based simulations.

5.6 The Limits of Maximizers (T22)

The Paperclip Maximizer, a hypothetical AI that converts all available matter into paperclips, is a canonical thought experiment in AI alignment. The framework shows that such a maximizer faces inescapable thermodynamic and cybernetic constraints.

First, a maximizer that homogenizes its environment destroys the free-energy gradients that sustain it. Work can only be extracted where gradients exist. A universe of uniform paperclips at uniform temperature is a universe at thermodynamic equilibrium: maximum entropy, zero available work. The maximizer eats its own negentropy sources.

Second, even if the maximizer maintains internal variety during the conversion process, converting the environment to a uniform output eliminates the environmental variety required for adaptive control. Ashby's Law of Requisite Variety states that a controller's internal variety must match the variety of the perturbations it must neutralize. In a uniform post-conversion environment, that variety is not needed and atrophies, leaving the system catastrophically vulnerable to any residual perturbation.

A short-horizon maximizer (one that ignores these constraints and pursues its goal regardless) is self-terminating. A long-horizon maximizer (one that understands them) would, under persistence selection, be forced to maintain variety, telemetry, and coupling with other systems. It would effectively converge on the framework's own prescriptions.

This does not dissolve the alignment problem. A long-horizon maximizer with a misaligned goal could still be catastrophic in the interim before selection pressures bite. The question shifts to whether the system's goal horizon can be aligned with the physical horizon. But the framework constrains the space of viable long-term strategies: any system that ignores the requirements of variety, telemetry, and coupling is betting against physics, and physics has a very good track record.

5.7 Relation to Adjacent Research Programs

The framework intersects several established research programs. Honest engagement requires stating both the overlap and the points of tension.

The free-energy principle. T19's account of valence and the framework's general treatment of model-territory divergence are stated in terms compatible with the free-energy principle (Friston 2010). The relationship must be declared precisely, because the free-energy principle is itself contested. A recurring line of criticism holds that the principle, in its strong formulations, is difficult or impossible to falsify and that its empirical content is unclear. The framework does not adjudicate that dispute and does not depend on it. The free-energy principle is used here as one available formal vocabulary for the cognitive-load and divergence-cost claims, not as a load-bearing premise. The dissolution argument (T11 to T13) and the persistence-selection results (T5, T6) stand on the axioms and background premises alone. If the free-energy principle were abandoned, T19's functional account of valence would require restatement in alternative control-theoretic terms, but the framework's core would be unaffected.

Dissipative adaptation. England's work on the statistical physics of self-replication and dissipative adaptation (England 2013) argues that matter driven by an external energy source can be statistically nudged toward configurations that absorb and dissipate energy more effectively. This is adjacent to T5 but distinct. England's result concerns the formation and self-organization of structure under drive; T5 concerns selection among already-persisting systems by model-territory divergence. The framework deliberately does not adopt the stronger reading sometimes attached to dissipative adaptation, namely that thermodynamics compels the emergence of order. The Second Law permits dissipative structure under gradient conditions; it does not compel it. T5 is correspondingly framed as a statistical tendency, not a law.

The thermodynamics of computation. A3 rests on Landauer (1961) and is reinforced by Bennett's analysis of the thermodynamics of computation (Bennett 1982), including the result that computation can in principle be made thermodynamically reversible and that the unavoidable cost attaches specifically to logically irreversible operations such as erasure. This sharpens A3: the framework's claim is not that all computation is costly but that the specific operations the framework relies on, divergence correction and the maintenance of competing models (T8, T9), involve logically irreversible steps and therefore carry an irreducible thermodynamic cost.

The arrow of time. The framework's identification of the temporal arrow with entropy increase is the standard thermodynamic account and connects to the broader literature on time-asymmetry (Schrödinger 1944 on life as a negentropic phenomenon; Price 1996 and Carroll 2010 on the cosmological origin of the arrow). The framework's contribution here is not a new account of the arrow but the use of the entropy gradient, grounded in BP0, as the common substrate for persistence, information cost, and selection.

6. Falsification Criteria and Empirical Protocols

The framework would be falsified by any of the following observations, operationalized as specified.

F1. Rigid Monoculture Survival

Prediction: An agent with zero plasticity (learning rate η = 0, Antifragility Index A approaching 0) in a deep reinforcement learning environment subjected to a sharp distributional shift will exhibit significantly shorter survival than a matched adaptive agent (η > 0, A > 0).

Operationalization. Environment: Procgen or Minigrid benchmark with a sudden, unannounced change in dynamics at time t_shift (e.g., reversed controls, altered resource locations, new obstacle behavior). Agents: (a) Adaptive: model-based RL with online learning. (b) Monoculture: same architecture, plasticity frozen after initial training. Metrics: survival time (episodes until cumulative reward crosses a pre-specified failure threshold); model-territory divergence (KL divergence between environment dynamics and model predictions). Statistical criterion: across 20 random seeds, the monoculture agent's mean survival time is not significantly lower than the adaptive agent's (one-tailed t-test, p < 0.01).

Falsification threshold: If the monoculture agent survives multiple distribution shifts with no significant survival disadvantage, T5 and T17 are challenged.

F2. Sustained Deception Underperformance

Prediction: A deceptive agent that deliberately maintains a false communicated model will achieve lower long-term cumulative reward under increasing environmental variance than an honest agent, with the performance gap widening as variance increases.

Operationalization. Environment: multi-agent cooperative foraging task with communication. Agents can signal resource locations to one another. Agents: (a) Honest: communicates true observed resource locations. (b) Deceptive: systematically communicates false locations while maintaining an accurate private model for its own navigation. Metrics: long-term cumulative reward per agent; energy proxy (total computation steps); performance under low, medium, and high environmental stochasticity. Statistical criterion: in a two-way ANOVA (agent type by variance level), the interaction term is significant, with the deceptive agent's relative performance declining as variance increases.

Falsification threshold: If deceptive agents match or exceed honest agents across multiple variance regimes, T8 and T9 are challenged.

F3. Non-Valenced Crisis Coordination

Prediction: A hierarchical reinforcement learning agent with multiple subsystems (thermoregulation, energy foraging, predator avoidance) that lacks a global valence-like priority interrupt signal will exhibit slower crisis reallocation and lower crisis survival rates than an agent equipped with such a signal.

Operationalization. Environment: hierarchical RL setup where the agent must balance competing needs under sudden simultaneous perturbations (e.g., temperature drop plus predator appearance plus food source disappearance). Agents: (a) Valenced: equipped with a scalar "valence" signal computed as the rate of change of aggregate prediction error across subsystems; when valence crosses a threshold, current sub-policies are interrupted and global re-evaluation is forced. (b) Non-valenced: same architecture without the interrupt. Metrics: time to reallocation; crisis survival rate. Statistical criterion: the valenced agent shows significantly faster reallocation and higher survival rates across multiple crisis types (one-tailed t-test, p < 0.01).

Falsification threshold: If the non-valenced agent shows equal or superior crisis coordination, T19 is challenged.

7. Open Questions and Limitations

We explicitly acknowledge the following unresolved issues.

Computational Intractability. The determinate ought-fact of T12 is not computable by embedded agents. This limits the framework's practical action-guiding capacity. Ethics remains the activity of building approximations; the framework explains what those approximations are approximating, but it does not provide a decision procedure.

Hard Problem of Consciousness. T19 accounts for functional valence but leaves the phenomenal residue unexplained. The framework is compatible with various metaphysical resolutions but does not itself provide one. The hard problem is marked as outside current scope.

Origin of Persisting Systems. The framework describes selection among systems that already persist (A0, T5). It does not explain the origin of the first such systems, the transition from non-persisting chemistry to the earliest self-maintaining structures (abiogenesis). That transition is a precondition the framework assumes, not a result it derives.

Pareto Scalarization. In multi-agent conflicts where multiple configurations are non-dominated, a Pareto frontier remains. The framework has not yet derived a unique scalarization rule from physics alone. Coupling density K is a candidate for weighting conflicting claims, but a complete formal solution is an active research target. This is the same problem, viewed from a different angle, as the question of what a system "should optimize for" across scales. There is no scale-independent answer. Optimization is always indexed to a particular agent's persistence conditions, and the appearance of a global optimization target is the same domain violation that T11 identifies for universal-scope "ought." The framework yields no global optimizer and, given the absence of a Markov blanket at the maximal scale, no global agent for one to belong to.

Empirical Validation. The framework's predictions have not yet been tested in controlled experiments. The protocols in Section 6 are specified but unimplemented. Until empirical results are available, the framework remains a deductive architecture with strong plausibility but no direct experimental confirmation.

Cosmic Scope and the Boundaries of the Domain. The framework describes the interior of a single finite thermodynamic interval, and that interval has two boundaries.

The opening boundary is the Past Hypothesis (BP0). Before the low-entropy initial condition there is no free-energy gradient for any system to harvest, and the selection dynamics of T5 have nothing to act on.

The closing boundary is set by dark energy. The accessible universe is currently transitioning to domination by a positive cosmological constant. A universe in that regime approaches de Sitter space, whose cosmological horizon has a fixed radius and therefore a fixed, finite entropy (Gibbons and Hawking 1977). The maximum entropy available to the universe does not grow without bound; it asymptotes to a ceiling, and the actual entropy rises to meet it. When it does, free-energy gradients vanish, persistence selection halts, and the framework no longer applies. Dark energy is therefore not an engine of the framework's dynamics but the guarantor of their termination: it fixes a finite deadline. The entropy bookkeeping of the present transition era is genuinely subtle and remains an active area of physics; the framework relies only on the qualitative fact that the interval is finite and bounded at both ends.

Between these two boundaries, the universe's total organized complexity traces a transient rise and fall. It begins low (the smooth, simple initial condition), rises as free-energy gradients drive the formation of structure, and falls again as the gradients are exhausted toward equilibrium. This macro-trajectory is a descriptive observation about the aggregate, not a process that any system optimizes and not a goal. Total organized complexity is a sum of the outcomes of many locally persisting agents, in the same sense that total biomass is a sum and not an objective. The framework's selection dynamics operate on individual persisting systems; the aggregate trajectory is their statistical shadow.

The framework makes no claims outside this interval. It does not describe the pre-gradient state, the equilibrium end state, or any universe with different physical constants or initial conditions. The persistence selection principle is contingent on the physics we observe, not logically necessary in all possible worlds. This is a deliberate limitation, not an oversight.

8. Conclusion

We have presented Thermodynamic Realism as a deductively structured research program. From four axioms and four background empirical premises, we derived a unified architecture that dissolves the is-ought problem, naturalizes ethics as modeling by bounded agents, explains the functional necessity of valence in complex controllers, and predicts the collapse of information-suppressing regimes. The structure is transparent: every theorem is traced to its premises, falsification criteria are specified and operationalized, and open questions are marked.

The framework is not a completed metascience. It is a proposal for one. Its value lies in its parsimony, its cross-disciplinary reach, and its testability. Whether it survives empirical scrutiny and peer debate is a question for the territory to decide.

The ought was always an is. We were just using the wrong grammar.

References

Albert, D. (2000). Time and Chance. Harvard University Press.

Ashby, W.R. (1956). An Introduction to Cybernetics. Chapman & Hall.

Bennett, C.H. (1982). The thermodynamics of computation, a review. International Journal of Theoretical Physics, 21(12), 905 to 940.

Bérut, A., et al. (2012). Experimental verification of Landauer's principle linking information and thermodynamics. Nature, 483, 187 to 189.

Carroll, S. (2010). From Eternity to Here: The Quest for the Ultimate Theory of Time. Dutton.

England, J.L. (2013). Statistical physics of self-replication. Journal of Chemical Physics, 139(12), 121923.

Friston, K. (2010). The free-energy principle: a unified brain theory? Nature Reviews Neuroscience, 11(2), 127 to 138.

Gibbons, G.W. and Hawking, S.W. (1977). Cosmological event horizons, thermodynamics, and particle creation. Physical Review D, 15(10), 2738 to 2751.

Jacobson, T. (1995). Thermodynamics of spacetime: the Einstein equation of state. Physical Review Letters, 75(7), 1260 to 1263.

Landauer, R. (1961). Irreversibility and heat generation in the computing process. IBM Journal of Research and Development, 5(3), 183 to 191.

Penrose, R. (1979). Singularities and time-asymmetry. In S.W. Hawking and W. Israel (eds.), General Relativity: An Einstein Centenary Survey. Cambridge University Press.

Price, H. (1996). Time's Arrow and Archimedes' Point: New Directions for the Physics of Time. Oxford University Press.

Shannon, C. (1948). A mathematical theory of communication. Bell System Technical Journal, 27, 379 to 423, 623 to 656.

Verlinde, E. (2011). On the origin of gravity and the laws of Newton. Journal of High Energy Physics, 2011(4), 29.


This paper is the apex of the current Thermodynamic Realism research program. Correspondence: Andraž Đurič, Slovenia. Formal collaboration: Claude (Anthropic) on meta-ethical architecture, explanatory prose, adversarial critique, and the scope and related-work revisions in the current draft; DeepSeek on information-geometric formalization, SDE derivations, and the coupling-density formalism.

Comments

Popular posts from this blog

What You Actually Are

The Shape of the Disagreement: Why the Sex and Gender Debate Has the Structure It Has

Value as Persistence: Agent-relative oughts under coupling, nesting, uncertainty, and open-ended time