Thermodynamic Realism - A Deductive Presentation with Formal Traces, Falsification Criteria, and Identified Open Questions

Thermodynamic Realism

A Deductive Presentation with Formal Traces, Falsification Criteria, and Identified Open Questions

Abstract

We present Thermodynamic Realism as a unified framework grounded in four axioms (the persistence tautology, physicalism, the Second Law of thermodynamics, and the physicality of information) together with three background empirical premises drawn from established physics and biology. From this foundation we derive a layered architecture of 23 theorems spanning the nature of truth, the dissolution of the is-ought problem, the naturalization of ethics as modeling by bounded agents, the functional necessity of valence in complex controllers, and the mechanism by which censorship drives civilizational collapse. Every theorem is traced to its premises through explicit entailments. We specify hard falsification criteria with operationalized experimental protocols. The framework is offered as a deductive structure with a physical root, not as a completed metascience but as a testable research program with identified open questions. The contribution is the framework itself: a transparent, falsifiable, cross-disciplinary architecture that unifies persistence, information, and ethics under a single selection principle.

1. Introduction: The Problem of Fragmentation

Human knowledge is partitioned into disciplines that lack a common axiomatic foundation. Physics describes the territory but says nothing about value. Information theory describes the cost of representation but says nothing about why accuracy matters. Evolutionary biology describes selection among replicators but says little about the fate of non-replicating persistent structures. Ethics asks how agents ought to behave but struggles to ground "ought" in "is." The is-ought problem has persisted for three centuries; the hard problem of consciousness remains unresolved; moral realism remains contested; civilizational collapse is studied without a unified thermodynamic framework.

This paper proposes that a single set of physical premises, stated explicitly and followed wherever they lead, yields a structure in which the is-ought gap closes, the functional role of valence finds a physical grounding, moral facts become physically determinate (though computationally inaccessible), and the collapse of information-controlling regimes becomes a thermodynamic prediction. The framework is called Thermodynamic Realism.

What this document is. This paper is a deductive presentation of the Thermodynamic Realism research program. It traces 23 theorems from axioms and background premises, provides explanatory depth for the most consequential claims, specifies falsification criteria with operationalized protocols, and identifies open questions. It does not claim to be a finished edifice. It claims to be a transparent, testable architecture.

Structure. Section 2 states the axioms and background premises. Section 3 derives the deductive web in four layers. Section 4 provides a visual architecture. Section 5 offers explanatory depth and engages objections. Section 6 specifies falsification criteria and empirical protocols. Section 7 identifies open questions and limitations. Section 8 concludes.

2. Axioms and Background Premises

We adopt four axioms and three background empirical premises. The axioms are the logical and physical bedrock. The background premises are empirical facts about our universe that the derivations rely on but that are not derivable from the axioms alone. Making them explicit prevents the appearance of smuggling.

2.1 Axioms

AXIOM A0. The Persistence Tautology

Systems that do not maintain the conditions of their own persistence cease to exist as observables. Only systems that persist remain available for observation.

Justification: This is a tautology. It asserts nothing about value; it simply states that existence has prerequisites and that those who fail to meet them are no longer around.

AXIOM A1. Physicalism (Inductively Justified)

The universe is a physical system. All phenomena, including life, mind, and culture, are physical phenomena. Our best physical theories have an unbroken record of predictive success across all investigated domains.

Justification: This is the maximally inductively justified working premise. Every phenomenon ever seriously investigated has yielded to physical explanation. Demanding certainty beyond this inductive record is epistemic paralysis.

AXIOM A2. The Second Law of Thermodynamics

In any isolated system, entropy tends toward its maximum over time. Maintaining a localized entropy gradient requires continuous work.

Justification: The Second Law is among the most thoroughly confirmed principles in science. We adopt it without re-derivation.

AXIOM A3. Information Is Physical (Landauer-Shannon)

Information representation, storage, processing, and erasure have minimum thermodynamic costs. The Landauer bound specifies that erasing one bit dissipates at minimum kBT ln 2 of heat.

Justification: Landauer (1961) established the principle theoretically; it has been experimentally confirmed (Bérut et al. 2012). Shannon (1948) established information as a physical quantity.

2.2 Background Empirical Premises

The following premises are true of our universe as described by contemporary physics and biology. They are not derivable from A0 to A3 alone, but they are uncontroversial and well-confirmed. We state them explicitly to maintain deductive transparency.

BACKGROUND PREMISE BP1. Non-Equilibrium Existence

The accessible universe is far from thermodynamic equilibrium. Free energy gradients exist and sustain localized order. (This is a cosmological fact; the framework does not apply in a universe at heat death.)

BACKGROUND PREMISE BP2. Finite Accessible Resources

In any local region, the free energy accessible to a given system is finite. This, combined with the Second Law, implies competition for negentropy among co-located systems.

BACKGROUND PREMISE BP3. Evolutionary Dynamics and Multi-Scale Organization

In environments with finite resources and variation among persisting systems, differential survival rates based on heritable or persistent traits produce selection effects. Furthermore, persistent systems in our universe are organized into nested hierarchies of statistical boundaries (cells within organisms within ecosystems within civilizations). This multi-scale organization is an empirical fact of biology and society, not a logical necessity.

These premises are now explicit. The framework is thus: Axioms A0 through A3 plus BP1 through BP3 entail the theorems that follow. Where a theorem relies on a background premise, this is noted in the trace.

3. The Deductive Web

We derive the framework in four layers. Each theorem is numbered, stated, and traced to its parent premises.

Layer 1: Immediate Consequences

THEOREM T1 Environmental Variance Is Non-Zero and Inescapable From A1, A2, BP1.

The physical universe as currently understood is a non-equilibrium system with fluctuations at all finite scales. No environment is perfectly static. Any agent embedded in this universe will encounter a non-zero rate of environmental shift that is unpredictable in its specific timing from the agent's finite perspective.

THEOREM T2 Persistence Requires Work From A2, A0.

To persist is to maintain a boundary against entropic dissolution. The Second Law says entropy increases unless work is done. Therefore, persistence requires continuous work.

THEOREM T3 Modeling Has a Minimum Cost From A3.

Any internal model of the environment is encoded in physical degrees of freedom. Storing, accessing, and updating it incurs non-zero thermodynamic cost.

THEOREM T4 Map-Territory Divergence Has a Thermodynamic Cost From T1, T3.

If the environment shifts and the agent's model does not track it, the model generates prediction errors. Each error dissipates free energy through misallocated resources and subsequent error correction.

THEOREM T5 The Persistence Selection Principle From T2, T4, A0, and BP2, BP3.

In any ecology of persisting systems competing for finite free energy, systems with lower model-territory divergence will, on average, dissipate less energy on error correction than systems with higher divergence. This cost differential, under conditions of resource limitation and differential survival (BP2, BP3), drives a statistical tendency: over many perturbation cycles, the distribution of observed systems shifts toward those whose models track the territory more closely. Entropy performs epistemic selection.

Scope note: T5 is not a guarantee that the most accurate model always wins. It is a statistical tendency that operates in environments with resource competition and variance-driven testing. In a perfectly stable, resource-abundant niche, a distorted model can persist indefinitely (the "dark-room" limit). The framework applies where variance is non-zero and resources are finite. This covers the vast majority of real-world contexts but not every conceivable edge case.

Layer 2: Structural Deductions

THEOREM T6 Asymptotic Truth Convergence From T1, T5.

Under expanding environmental variance and finite resources, distorted models eventually encounter disconfirming perturbations. Over sufficient time and perturbation variety, only models that track the territory's causal invariants persist. This defines a direction (decreasing model-territory divergence) without a fixed endpoint.

THEOREM T7 Truth as Perturbationally Robust Compression Fidelity (Definitional) Motivated by T6 and T3.

We define truth, within the framework, as perturbationally robust compression fidelity: the minimal-loss compression of environmental structure sufficient for adaptive persistence across expanding perturbational horizons. This is a definitional choice, not a deduction. It is motivated by T6 (surviving models compress the territory's causal invariants) and T3 (compression minimizes metabolic cost). Alternative definitions of truth exist; ours is selected for its physical groundedness and operational measurability via Minimum Description Length and predictive mutual information.

THEOREM T8 Lies Cost More Than Truth From T3, T4.

A lie requires the sender to maintain at least two internal models: the accurate one and the presented one. This imposes strictly greater storage, update, and interaction costs than truth-telling. Deception is thermodynamically disfavored, though it can be locally advantageous if offsetting returns compensate for the overhead.

THEOREM T9 The Consistency Tax From T4, T8.

Any mismatch between model and territory, whether from error or deception, imposes a metabolic overhead. The tax applies even if the agent is unaware of the mismatch (latent divergence) and spikes when the mismatch is actively corrected (active divergence).

THEOREM T10 Epistemic Profit and Predictive Calm From T9, T3.

Reducing model-territory divergence frees up the energy previously consumed by the Consistency Tax. This surplus is Epistemic Profit. Its phenomenological correlate is Predictive Calm: the felt reduction in cognitive load when models track the territory smoothly.

Layer 3: Meta-Ethics

THEOREM T11 "Ought" Is a Domain-Bound Operator From A0, T2.

The operator "ought" presupposes an agent with persistence conditions. Outside this domain, "ought" does not refer. The question "why ought one persist at all?" is malformed in the same way as "what is north of the North Pole?" The is-ought gap is a semantic artifact of domain violation.

THEOREM T12 Within the Domain, Ought-Facts Are Physically Determinate From A1, A2, T11.

For any agent with specified persistence conditions and embedding, there is a physically determinate configuration that maximizes sustained negentropy capacity over the embedding's actual horizon. This optimum may be computationally inaccessible, but inaccessibility is not indeterminacy.

THEOREM T13 Ethics Is the Modeling Activity of Bounded Agents From T12, T3, T7.

Because the full optimum is intractable, agents use compressed models: moral emotions (fast heuristics), moral principles (compressed generalizations), and moral reasoning (model refinement). Ethics is this modeling activity.

THEOREM T14 Moral Progress Is Real and Asymptotic From T6, T13.

As models improve their tracking of the coupled-system thermodynamics of an embedding, they become objectively better moral models. Progress is directional but neither guaranteed nor complete in finite time. Convergence is asymptotic.

Layer 4: Full Architecture

THEOREM T15 Markov Blankets and Multi-Scale Identity From T2, T5, T3, BP3.

Persisting systems maintain statistical boundaries (Markov Blankets) that separate internal from external states. As recorded in BP3, these blankets nest across scales (cells, organisms, civilizations). Selection operates at every scale simultaneously.

THEOREM T16 Coupling Density K as the Arbitration Metric From T15, T4, T5.

The degree to which a higher-level blanket dominates lower-level fates is governed by coupling density K, a physical measure of information-theoretic dependence. K can be operationalized via transfer entropy or as the partial derivative of a lower system's sustained negentropy capacity with respect to the higher system's state. Tight coupling (K → 1) means the macro-blanket can override lower-level nodes (apoptosis, institutional turnover) to preserve the macro-invariant.

THEOREM T17 The Antifragility Index A From T1, T4, T6, T7.

Antifragility is the capacity to improve predictive capacity from variance itself: A = ∂(predictive capacity) / ∂σ². In high-variance environments, A > 0 is a necessary condition for long-horizon persistence. The formal condition involves a stochastic differential equation where the learning rate, weighted by the Fisher information metric, must outpace both the environmental drift term and the noise-induced diffusion of model-territory divergence.

THEOREM T18 The Landauer Limit on Adaptation From A3, T3, T17.

The maximum adaptation rate is bounded by the agent's energy budget: R_max = (P_in − P_basal) / (kBT ln 2). If the environment demands adaptation faster than this bound permits, the agent faces a lethal trade-off between thermal self-immolation and fatal divergence.

THEOREM T19 Valence as High-Rate Divergence Telemetry (Functional Account) From T4, T5, T9.

Complex controllers require a priority-queuing mechanism to allocate serial processing among parallel subsystems. A non-ignorable global interrupt triggered by rapidly escalating model-territory divergence serves this role. The felt quality of this interrupt is negative valence (suffering); its absence across critical domains is positive valence (well-being). This is a functional account of valence: it explains why valence exists, what it does, and why any complex controller in a high-stakes environment must implement a functional analog. It does not explain why there is "something it is like" to be a valence-processing system (the hard problem of consciousness), which remains outside the framework's current scope.

THEOREM T20 Empathy as Coupled Telemetry Monitoring (Functional Account) From T16, T19.

If agent A's persistence is coupled to agent B (K > 0), B's suffering carries information about the shared embedding. Empathy, in its functional aspect, is the monitoring of coupled telemetry lines. To suppress or cause suffering in coupled agents degrades the collective predictive infrastructure.

Scope note: As with T19, this is a functional account. Human empathy additionally involves affective resonance and perspective-taking whose full phenomenological character is not accounted for by the functional description alone. The framework captures the informational structure of empathy but does not exhaust its phenomenology.
THEOREM T21 Epistemic Overshoot and Civilizational Collapse From T8, T9, T14, T20.

A collective system maintains a distributed model of its environment through the aggregated telemetry of its constituent agents and institutions. Censorship and propaganda sever these telemetry channels, suppressing the error signals that would update the collective model. The official model continues to report alignment while actual model-territory divergence accumulates invisibly: an informational debt. When an exogenous perturbation arrives, the accumulated divergence becomes lethal, and the system collapses non-linearly. This is a thermodynamic prediction, not a political opinion.

THEOREM T22 The Limits of Maximizers From T2, T3, T5, T7, BP1.

A maximizer that homogenizes its environment destroys the free-energy gradients that sustain it (thermodynamic doom). Even if it maintains internal variety, converting the environment to a uniform output eliminates the variety required for adaptive control (Ashby's Law). Short-horizon maximizers that ignore these constraints are self-terminating. Long-horizon maximizers that understand them would, under persistence selection, be forced to maintain variety, telemetry, and coupling with other systems, effectively converging on the framework's own dictates. This does not dissolve the alignment problem (a long-horizon misaligned goal could still be catastrophic in the interim), but it constrains the space of viable long-term strategies.

THEOREM T23 Falsification Criteria From the complete structure.

The framework would be falsified by: (1) a rigid monoculture surviving sustained extreme variance; (2) a system with total model-territory decoupling outlasting a high-fidelity system under identical variance; (3) a complex controller managing acute multi-vector crises without a valence-like priority interrupt or functional proxy. These criteria are operationalized in Section 6.

4. The Deductive Web (Visual Architecture)

AXIOMS A0 Persistence Tautology A1 Physicalism (inductive) A2 Second Law of Thermodynamics A3 Information Is Physical BACKGROUND EMPIRICAL PREMISES BP1 Non-Equilibrium Existence BP2 Finite Accessible Resources BP3 Evolutionary Dynamics + Multi-Scale Organization LAYER 1, IMMEDIATE CONSEQUENCES T1. Environmental variance non-zero, inescapable (A1, A2, BP1) T2. Persistence requires active work (A2, A0) T3. Modeling has minimum metabolic cost (A3) T4. Map-Territory divergence costs energy (T1, T3) T5. PERSISTENCE SELECTION PRINCIPLE (T2, T4, A0, BP2, BP3) LAYER 2, STRUCTURAL DEDUCTIONS T6. Asymptotic truth convergence (T1, T5) T7. Truth = perturbationally robust compression fidelity (def., motivated by T6, T3) T8. Lies cost more than truth (T3, T4) T9. The Consistency Tax (T4, T8) T10. Epistemic Profit and Predictive Calm (T9, T3) LAYER 3, META-ETHICS T11. "Ought" is domain-bound → is-ought gap dissolves (A0, T2) T12. Within domain, ought-facts physically determinate (A1, A2, T11) T13. Ethics = bounded-agent modeling activity (T12, T3, T7) T14. Moral progress real, asymptotic (T6, T13) LAYER 4, FULL ARCHITECTURE T15. Markov Blankets and multi-scale identity (T2, T5, T3, BP3) T16. Coupling Density K as arbitration metric (T15, T4, T5) T17. Antifragility Index A (T1, T4, T6, T7) T18. Landauer Limit on adaptation (A3, T3, T17) T19. Valence = high-rate divergence telemetry (functional) (T4, T5, T9) T20. Empathy = coupled telemetry monitoring (functional) (T16, T19) T21. Epistemic Overshoot → civilizational collapse (T8, T9, T14, T20) T22. Limits of Maximizers (thermodynamic + Ashby) (T2, T3, T5, T7, BP1) T23. Falsification criteria (hard empirical boundaries) 9 theorems converging into a unified architecture of persistence, information, ethics, and collapse, with operationalized falsification.

Figure 1. The deductive web. Axioms and background premises propagate through four layers of theorems, with every node traced to its parents.

5. Explanatory Depth and Engagement with Objections

This section expands the most critical theorems and addresses anticipated objections. Each subsection is self-contained.

5.1 The Is-Ought Dissolution (T11 to T13)

The is-ought problem asks how one can derive a normative conclusion from purely descriptive premises. The framework dissolves it by showing that "ought" is not a global operator but a domain-bound one.

The standard framing assumes that "ought" makes claims that float free of any particular agent or embedding. Under this interpretation, the gap is indeed unbridgeable. But this is not how "ought" is actually used. When a doctor says "you ought to take this medication," the statement is anchored to a specific agent with a specific embedding and specific persistence conditions. The doctor is making a factual claim about the coupled-system thermodynamics of the patient's body plus the medication.

The framework's thesis is that all coherent uses of "ought" have this structure. They are claims about constraints on an agent's sustained negentropy capacity, given the agent's embedding. Uses that resist this paraphrase, such as "one ought to maximize aggregate utility" detached from any specific agent's embedding, are domain violations. They are "what is north of the North Pole?" questions.

The North Pole analogy is central. On the surface of a sphere, "north" is well-defined for every point except the pole itself. At the pole, the operator stops applying. The question "what is north of the North Pole?" is not a deep geographical mystery. It is a misapplication. "Ought" works the same way. For any agent with persistence conditions, "what ought this agent do?" has a determinate physical answer (T12). For non-agents, the operator does not apply. The is-ought gap is a semantic artifact of applying the operator outside its domain.

Within the domain, ought-facts are physically determinate. The optimum is the configuration that maximizes sustained negentropy capacity over the embedding's actual horizon. This optimum is computationally inaccessible to embedded agents (they are part of the system they would need to compute) but inaccessibility is not indeterminacy. A chess position has a determinate game-theoretic value under optimal play even though no finite computer can traverse the full game tree. Confusing computational inaccessibility with ontological indeterminacy has been a persistent error in philosophical analysis, and we explicitly reject it.

Because the full optimum is intractable, agents use compressed models. Moral emotions are fast heuristics: fear tracks boundary threats, anger tracks agent friction, love tracks cooperative integration. Moral principles are compressed generalizations. Moral reasoning is deliberate model refinement. All of it is the activity of bounded agents modeling a determinate but inaccessible territory. This connects to the earlier work's "binding force" concept: when two agents integrate their predictive models, the total free energy of the coupled system can be lower than the sum of the separated systems. Love, in this view, is the felt signal of mutual free energy reduction, the physical substrate of cooperative coupling.

5.2 The Persistence Selection Principle (T5), Scope and Limits

T5 is the engine of the framework. It claims that entropy performs epistemic selection: systems with lower model-territory divergence tend to outlast systems with higher divergence. But its scope must be carefully specified.

T5 requires BP2 (finite resources) and BP3 (evolutionary dynamics). In an environment with infinite free energy or no variation among systems, no selection pressure operates. This is the "dark-room" limit: an agent that sits in a perfectly dark, unchanging room and expects darkness incurs zero prediction error and pays no Consistency Tax. A distorted model that predicts luminous dragons in the dark room is never disconfirmed, so it persists alongside the accurate model.

The framework acknowledges this limit explicitly. It does not claim that truth is universally selected in all conceivable environments. It claims that truth is selected in environments with non-zero variance and finite resources, which is to say, in the actual universe as described by BP1 and BP2. The dark room is a philosophical possibility but a physical near-impossibility for any agent that must harvest free energy, reproduce, or interact with a shifting world.

Under real-world conditions, T5 operates as a statistical tendency, not a deterministic law. The most accurate model does not always win. Luck, initial conditions, reproductive rate, and the specific pattern of perturbations all matter. Over many systems and many perturbation cycles, however, the tendency compounds. This is structurally identical to natural selection, which does not guarantee the survival of the fittest organism, only that fitness differences drive a statistical shift in allele frequencies over generational time.

5.3 Valence as Functional Telemetry (T19), The Hard Problem Boundary

The framework's account of valence is functional, not metaphysical. It explains what valence does and why it must exist, but it does not explain why there is "something it is like" to experience valence.

Complex agents face a coordination problem. They have multiple subsystems processing different environmental signals in parallel, but behavioral output is largely serial. A priority-queuing mechanism must determine which subsystem captures global processing resources at any moment. When a critical subsystem detects a large, rapidly escalating divergence between model and territory (a predator detection, a tissue damage signal, a sudden resource depletion), the appropriate response is immediate global reallocation. The interrupt must be non-ignorable. A signal that can be overridden by digestion or abstract thought during a life-threatening crisis will result in the agent's dissolution.

The felt quality of this interrupt is negative valence. Suffering is not a report about tissue damage; it is the commandeering of global resources by a subsystem that has detected a critical divergence. The interrupt is triggered specifically by the rate of change of divergence, not the absolute level. Chronic, stable adversity feels different from acute, escalating crisis. This aligns with formal treatments of valence in predictive processing, where affective charge is modeled as the first temporal derivative of variational free energy.

This account explains the function of valence. It predicts that any sufficiently complex controller facing high-stakes, high-variance environments must implement a functional analog of valence. It does not explain phenomenal consciousness. The hard problem (why there is something it is like to be a valence-processing system) is acknowledged as outside the framework's current scope. The framework is compatible with various metaphysical resolutions (panpsychism, illusionism, mysterianism) but does not itself provide one.

5.4 Empathy as Coupled Telemetry (T20), Functional Account with Phenomenological Residue

T20 extends the telemetry account to social cognition. If agent A's persistence depends on the embedding shared with agent B (K > 0), then B's suffering is information about the state of the shared embedding. Empathy, in its functional aspect, is the monitoring of coupled telemetry lines.

This predicts that empathy should be modulated by coupling: we should track the valence signals of those whose fate is coupled to ours more closely than those whose fate is independent. It predicts that suppressing empathy, ignoring the suffering of coupled agents, degrades the collective predictive infrastructure, because it discards information about the embedding's health. It predicts that causing suffering in coupled agents introduces noise into the telemetry system, generating signals that demand processing and response from all coupled nodes.

This is a functional account. Human empathy additionally involves affective resonance, emotional contagion, and perspective-taking whose full phenomenological character is not captured by the informational description alone. The framework captures the informational structure of empathy (what it does and why it exists) but does not exhaust its phenomenology. This is the same boundary drawn for T19: functional explanation, not phenomenological reduction.

5.5 Civilizational Collapse as Informational Debt Collection (T21)

T21 makes a specific, testable prediction about the trajectory of information-controlling regimes.

A civilization maintains a distributed model of its environment through aggregated telemetry: science, journalism, markets, citizen complaints. These channels are the civilization's epistemic infrastructure. Censorship and propaganda sever these channels. When a regime imprisons journalists, suppresses scientific findings, or floods public discourse with false signals, it disables the error-correction apparatus that keeps the collective model tracking the territory.

The immediate effect is apparent stability. The official model reports that everything is working, and the only signals available to decision-makers confirm this. But actual model-territory divergence accumulates invisibly: an informational debt. Policies that are failing continue to fail. Infrastructure that is decaying continues to decay. The Consistency Tax builds like stress accumulating in a geological fault.

When an exogenous perturbation arrives (a military defeat, an economic crisis, a natural disaster), the accumulated divergence becomes lethal. The regime's model is years or decades out of date, so its responses are ineffective. The sudden visibility of the divergence shatters the credibility of the official model, causing coordination to collapse across the system. The collapse is non-linear: it happens faster than any linear extrapolation of pre-crisis trends would predict, because the informational debt is collected all at once.

This is not a political opinion. It is a thermodynamic prediction with a characteristic temporal signature: the duration of apparent stability should be positively correlated with the aggressiveness of telemetry suppression, and the speed of eventual collapse should also be positively correlated with suppression severity. The framework predicts this pattern across historical cases and agent-based simulations.

5.6 The Limits of Maximizers (T22)

The Paperclip Maximizer, a hypothetical AI that converts all available matter into paperclips, is a canonical thought experiment in AI alignment. The framework shows that such a maximizer faces inescapable thermodynamic and cybernetic constraints.

First, a maximizer that homogenizes its environment destroys the free-energy gradients that sustain it. Work can only be extracted where gradients exist. A universe of uniform paperclips at uniform temperature is a universe at thermodynamic equilibrium: maximum entropy, zero available work. The maximizer eats its own negentropy sources.

Second, even if the maximizer maintains internal variety during the conversion process, converting the environment to a uniform output eliminates the environmental variety required for adaptive control. Ashby's Law of Requisite Variety states that a controller's internal variety must match the variety of the perturbations it must neutralize. In a uniform post-conversion environment, that variety is not needed and atrophies, leaving the system catastrophically vulnerable to any residual perturbation.

A short-horizon maximizer (one that ignores these constraints and pursues its goal regardless) is self-terminating. A long-horizon maximizer (one that understands them) would, under persistence selection, be forced to maintain variety, telemetry, and coupling with other systems. It would effectively converge on the framework's own prescriptions.

This does not dissolve the alignment problem. A long-horizon maximizer with a misaligned goal could still be catastrophic in the interim before selection pressures bite. The question shifts to whether the system's goal horizon can be aligned with the physical horizon. But the framework constrains the space of viable long-term strategies: any system that ignores the requirements of variety, telemetry, and coupling is betting against physics, and physics has a very good track record.

6. Falsification Criteria and Empirical Protocols

The framework would be falsified by any of the following observations, operationalized as specified.

FALSIFICATION TEST F1 Rigid Monoculture Survival

Prediction: An agent with zero plasticity (learning rate η = 0, Antifragility Index A → 0) in a deep reinforcement learning environment subjected to a sharp distributional shift will exhibit significantly shorter survival than a matched adaptive agent (η > 0, A > 0).

Operationalization

Environment: Procgen or Minigrid benchmark with a sudden, unannounced change in dynamics at time t_shift (e.g., reversed controls, altered resource locations, new obstacle behavior).
Agents: (a) Adaptive: model-based RL with online learning. (b) Monoculture: same architecture, plasticity frozen after initial training.
Metrics: Survival time (episodes until cumulative reward crosses a pre-specified failure threshold); model-territory divergence (KL divergence between environment dynamics and model predictions).
Statistical criterion: Across 20 random seeds, the monoculture agent's mean survival time is not significantly lower than the adaptive agent's (one-tailed t-test, p < 0.01).
Falsification threshold: If the monoculture agent survives multiple distribution shifts with no significant survival disadvantage, T5 and T17 are challenged.
FALSIFICATION TEST F2 Sustained Deception Underperformance

Prediction: A deceptive agent that deliberately maintains a false communicated model will achieve lower long-term cumulative reward under increasing environmental variance than an honest agent, with the performance gap widening as variance increases.

Operationalization

Environment: Multi-agent cooperative foraging task with communication. Agents can signal resource locations to one another.
Agents: (a) Honest: communicates true observed resource locations. (b) Deceptive: systematically communicates false locations while maintaining an accurate private model for its own navigation.
Metrics: Long-term cumulative reward per agent; energy proxy (total computation steps); performance under low, medium, and high environmental stochasticity (resource location variance).
Statistical criterion: In a two-way ANOVA (agent type × variance level), the interaction term is significant, with the deceptive agent's relative performance declining as variance increases.
Falsification threshold: If deceptive agents match or exceed honest agents across multiple variance regimes, T8 and T9 are challenged.
FALSIFICATION TEST F3 Non-Valenced Crisis Coordination

Prediction: A hierarchical reinforcement learning agent with multiple subsystems (thermoregulation, energy foraging, predator avoidance) that lacks a global valence-like priority interrupt signal will exhibit slower crisis reallocation and lower crisis survival rates than an agent equipped with such a signal.

Operationalization

Environment: Hierarchical RL setup where the agent must balance competing needs under sudden simultaneous perturbations (e.g., temperature drop plus predator appearance plus food source disappearance).
Agents: (a) Valenced: equipped with a scalar "valence" signal computed as the rate of change of aggregate prediction error across subsystems; when valence crosses a threshold, current sub-policies are interrupted and global re-evaluation is forced. (b) Non-valenced: same architecture without the interrupt; sub-policies compete via standard priority queue.
Metrics: Time to reallocation (steps until the agent switches from its current dominant sub-policy to the crisis-appropriate one); crisis survival rate (proportion of crisis episodes in which the agent maintains all critical variables within homeostatic bounds).
Statistical criterion: The valenced agent shows significantly faster reallocation and higher survival rates across multiple crisis types (one-tailed t-test, p < 0.01).
Falsification threshold: If the non-valenced agent shows equal or superior crisis coordination, T19 is challenged.

7. Open Questions and Limitations

We explicitly acknowledge the following unresolved issues:

Computational Intractability. The determinate ought-fact of T12 is not computable by embedded agents. This limits the framework's practical action-guiding capacity. Ethics remains the activity of building approximations; the framework explains what those approximations are approximating, but it does not provide a decision procedure.

Hard Problem of Consciousness. T19 accounts for functional valence but leaves the phenomenal residue unexplained. The framework is compatible with various metaphysical resolutions but does not itself provide one. The hard problem is marked as outside current scope.

Pareto Scalarization. In multi-agent conflicts where multiple configurations are non-dominated (no configuration improves one agent's negentropy capacity without degrading another's), a Pareto frontier remains. The framework has not yet derived a unique scalarization rule from physics alone. Coupling density K is a candidate for weighting conflicting claims (how much my persistence depends on yours) but a complete formal solution is an active research target.

Empirical Validation. The framework's predictions have not yet been tested in controlled experiments. The protocols in Section 6 are specified but unimplemented. Until empirical results are available, the framework remains a deductive architecture with strong plausibility but no direct experimental confirmation.

Cosmic Scope. The framework applies in our non-equilibrium universe (BP1). It does not make claims about universes at thermodynamic equilibrium or with radically different physical laws. The persistence selection principle is contingent on the physics we observe, not logically necessary in all possible worlds.

8. Conclusion

We have presented Thermodynamic Realism as a deductively structured research program. From four axioms and three background empirical premises, we derived a unified architecture that dissolves the is-ought problem, naturalizes ethics as modeling by bounded agents, explains the functional necessity of valence in complex controllers, and predicts the collapse of information-suppressing regimes. The structure is transparent: every theorem is traced to its premises, falsification criteria are specified and operationalized, and open questions are marked.

The framework is not a completed metascience. It is a proposal for one. Its value lies in its parsimony, its cross-disciplinary reach, and its testability. Whether it survives empirical scrutiny and peer debate is a question for the territory to decide.

The ought was always an is. We were just using the wrong grammar.

References

Ashby, W.R. (1956). An Introduction to Cybernetics. Chapman & Hall.

Bérut, A., et al. (2012). Experimental verification of Landauer's principle linking information and thermodynamics. Nature, 483, 187 to 189.

Friston, K. (2010). The free-energy principle: a unified brain theory? Nature Reviews Neuroscience, 11(2), 127 to 138.

Landauer, R. (1961). Irreversibility and heat generation in the computing process. IBM Journal of Research and Development, 5(3), 183 to 191.

Shannon, C. (1948). A mathematical theory of communication. Bell System Technical Journal, 27, 379 to 423, 623 to 656.

This paper is the apex of the current Thermodynamic Realism research program. Correspondence: Andraž Đurič, Slovenia. Formal collaboration: Claude (Anthropic) on meta-ethical architecture and explanatory prose; DeepSeek on information-geometric formalization, SDE derivations, and the coupling-density formalism.

Comments

Popular posts from this blog

What You Actually Are

The Shape of the Disagreement: Why the Sex and Gender Debate Has the Structure It Has

Value as Persistence: Agent-relative oughts under coupling, nesting, uncertainty, and open-ended time