The Shape of an Idea

The Shape of an Idea

How a model returns the thought you could not put into words — and the test that separates faithful rendering from plausible confabulation

Analysis

Andraž Đurič

Version 1 · June 2026

You have a thought you cannot get out cleanly. You write it down badly — half-formed, hedged, abandoned mid-sentence, more noise than signal — and hand it to a language model. What comes back is closer to what you meant than what you wrote. It feels as though the model read your mind.

It did not. What happened is more ordinary, and worth understanding, because the same process that produces a faithful rendering also produces a convincing forgery, and from the inside the two are indistinguishable. The short version: the model decompressed a signal that carried more constraint than it appeared to; whether the result is your idea or the model's turns on a single thing — whether you hold an internal model to check it against; the check is recognition, and recognition can be faked; the guard against the fake is committing to the shape before you see the output.

What is claimed here, and at what strength. Three things are established and borrowed: that people hold knowledge they cannot fully put into words (Polanyi, as the nearest anchor — the case here is narrower, an idea not yet serialized rather than knowledge resistant to articulation in principle); that a generic, flattering description is readily accepted as an accurate portrait of oneself (the Forer effect); and that shaping an account to fit a result after seeing it corrupts the test (HARKing). One thing is conjectural, not established: the mechanism — that the model performs constrained interpolation on a signal whose redundancy pins the target — is one hypothesis about what the process is doing, stated as such and revisable, not a measured fact about how these systems work, and very possibly only part of it; the match may well be multifactorial, with constrained decompression one contributor among several. What this piece adds is narrow: the distinction between two regimes, keyed to whether an internal model is available to verify against, and the use of pre-registration as the line between trustworthy and spoofable acceptance. The component observations are not new; the application, and the line drawn through them, are the contribution.

The signal does the selecting

It is tempting to credit the model's breadth — the vast range of text it has compressed — for the quality of what comes back. Breadth is necessary, but it explains the wrong thing. Breadth supplies a rich space of well-formed expressions to land on; it is the menu. It does not explain why the output lands on your meaning rather than on some generically good expression pointed elsewhere. That selection is done by your signal. A message that is malformed on the surface can still encode the relations, the direction, and the tensions you have already anticipated, and those constraints pin the target to a narrow region. The model returns the nearest well-formed serialization consistent with them. Call it constrained decompression: the apparent noise is not the whole signal, and what looks like the model supplying the idea is mostly the model satisfying constraints you supplied.

This account of the mechanism is a hypothesis, not an established fact, and is stated as one — one candidate explanation, very possibly partial; the actual match may be multifactorial, with constrained decompression at most a part of what produces it. It is the kind of claim that should be gradeable and given up if it stops fitting. What supports it at the premise level is narrow and solid: natural language is heavily redundant, so a partial, broken expression still carries most of the structure that fixes its meaning — Shannon's measurements of printed English are the standard demonstration that text is far more constrained than its surface suggests. The model does real work on top of that — finding the conventional form, connecting your shape to positions already named in some literature, surfacing the standard expression you were reaching for — but the thing that makes the output yours is the constraint you put in, not the breadth it drew on. Breadth determines the quality of the available expressions; your signal determines which one is selected.

Two regimes

The same machine produces two outcomes that feel identical from the inside.

Regime one. You hold the understood shape and lack only the serialization. The model decompresses it, and you verify the result by recognition against your internal template. The rendering is faithful — and it is trustworthy because you can check it, not because the model is reliable. The verification is yours; the model only proposed.

Regime two. You hold only a gesture, with no determinate internal shape. The model supplies a plausible completion. You cannot tell the difference, because there is no template to check against. The output is the model's shape wearing your question's clothes, and it may be coherent, even good — but it is not a rendering of anything you held.

What separates the two is whether there is an internal model to verify against. That is the whole distinction, and it dissolves a worry that otherwise looks serious: that using the model launders its idea into yours. In regime one you are not receiving an idea, you are recognizing a rendering — acceptance is active matching against a held template, not passive reception. In regime two you are receiving an idea and cannot tell. The hazard is precisely that the two feel the same.

Whether the shape is yours does not turn on where the words came from. The words are always borrowed — from the literature the model compressed, and behind that from other people. Provenance and ownership are orthogonal: the shape is yours if it matches what you understood, and judging it by its source — "the model wrote it, so it is the model's," or "I prompted it, so it is mine" — is the genetic fallacy run in one direction or the other. And every component can be borrowed while the configuration is new. Novelty lives in the arrangement, not the parts.

Recognition can be faked

Regime one rests on recognition, and recognition is spoofable. A serialization that is plausible and slightly flattering is accepted as what you meant even when it was not — and, worse, is afterward remembered as your prior intent. This is the Forer effect generalized: people rate generic descriptions as accurate accounts of themselves, the more readily when the description is favorable. Turned on this setting, a good completion can be installed retroactively as the thing you were trying to say, with the installation experienced as recall.

So "yes, exactly" is weak evidence. It is the response a good forgery is built to produce, and it arrives with the same felt certainty whether the rendering is faithful or fitted.

The guard is pre-registration. If you can state the shape's commitments before seeing the output — it must do this, must not collapse into that, must resolve the tension between these two — and the output satisfies those stated commitments, that is far stronger than recognition after the fact. A commitment fixed in advance cannot be quietly bent to fit what came back. Recognition after the fact is the cognitive form of HARKing — hypothesizing after the results are known — where the standard for "what I meant" is adjusted, unnoticed, to match the output. Stating the constraints first is what keeps that adjustment from happening, because there is a fixed record to fail against.

This is why listing your objections and counter-moves before you read the rendering is not throat-clearing. It is the pre-registration that makes the rendering checkable: a target pinned before serialization, against which the output either does or does not conform.

The generative case

There is a third outcome, neither faithful rendering nor forgery. Sometimes the act of serialization surfaces an implication you had not seen, and the shape itself changes. The output is no longer a rendering of what you held; it is an extension of it.

This is the most valuable outcome for real work and the one that most needs vigilance, and for the same reason: the extension is the model's contribution, and it carries the model's errors and biases into your thinking. The earlier guard does not reach it. In regime one you check the rendering against your template; in the generative case there is, by definition, no prior template for the new part, because the new part is new. So it has to be verified the hard way — against the world, or against argument — and not against your prior understanding. The feeling of "that is right" carries no weight here, because there was nothing to recognize. The generative case is where the model adds the most and where recognition is worth the least.

What this is and is not

This is not an account of how language models work; the mechanism is a reconstruction, marked as one. It is not a claim that the model understands anything; whatever the model does, the verification is done by you. It is not advice to distrust the output — in regime one, with the commitments pre-registered, the output is trustworthy, and the entire point is to know which regime you are in rather than to fix a posture in advance.

What it is: a way to tell a faithful rendering from a plausible confabulation when the two feel the same from the inside, and a procedure that follows from the distinction — hold an internal model, state its commitments before you look, verify by satisfaction rather than by recognition, and treat anything genuinely new as owed a separate check. The rendering you accept is yours when you held the shape and it passed; it is the model's, wearing your clothes, when you did not and could not tell.

Standing of this document

An analysis, not a result, and its parts do not share a status. The two borrowed psychological findings — that generic flattering descriptions are accepted as self-accurate (Forer), and that fitting an account to a result after the fact corrupts it (HARKing) — carry the confidence they carry in their own literatures; nothing here strengthens or weakens them. The premise that a broken signal still carries most of its meaning rests on the redundancy of natural language, which is measured and secure; the mechanism built on it — constrained decompression — is a hypothesis about what the process is doing, not a finding about these systems, offered as one candidate and probably only a partial one, since the match may be multifactorial and the full causal story is not known. Nothing downstream depends on it: the contribution — the two-regime distinction keyed to verifiability, and pre-registration as the line between trustworthy and spoofable acceptance — needs only that the signal carries constraint a holder of the shape can verify against, which holds whatever the mechanism turns out to be. That part is conceptual, and stands or falls on the argument rather than on data. The kill condition: exhibit a case in which you held no internal model and could state no commitment in advance, yet the model's output was reliably your meaning rather than a plausible completion you accepted as such — and the regime distinction fails. None has been offered here. Corrections and counterexamples are welcome and change the document.

References

  1. Polanyi, M. (1966). The Tacit Dimension. Doubleday, Garden City, NY. Knowledge held and used but not fully statable — "we know more than we can tell"; the nearest anchor for an idea present before it is articulated, though the phenomenon here is narrower: a shape not yet serialized, rather than knowledge resistant to articulation in principle. Verified
  2. Shannon, C. E. (1951). "Prediction and Entropy of Printed English." Bell System Technical Journal, 30(1), 50–64. The redundancy of natural language: text is far more predictable, and so far more constrained, than its surface suggests. Supports only the premise that a partial, malformed expression still carries most of what fixes its meaning — not the full mechanism built on that premise. Verified
  3. Forer, B. R. (1949). "The Fallacy of Personal Validation: A Classroom Demonstration of Gullibility." The Journal of Abnormal and Social Psychology, 44(1), 118–123. Generic descriptions are accepted as accurate portraits of oneself, especially when favorable; the spoofability of recognition, and the basis for retroactive installation of an output as prior intent. Verified
  4. Kerr, N. L. (1998). "HARKing: Hypothesizing After the Results Are Known." Personality and Social Psychology Review, 2(3), 196–217. Adjusting a hypothesis to fit a result after seeing it; the cognitive form of post-hoc recognition, and the failure that pre-registration is designed to prevent. Verified

Author: Andraž Đurič, Slovenia. Written in dialogue with Claude (Anthropic), a credited collaborator, following exchanges with other language models. Contributions are judged on content rather than origin; doing otherwise would be the genetic fallacy. Text licensed CC BY 4.0.

Comments

Popular posts from this blog

What You Actually Are

The Shape of the Disagreement: Why the Sex and Gender Debate Has the Structure It Has

Value as Persistence: Agent-relative oughts under coupling, nesting, uncertainty, and open-ended time