Everything going on in AI - updated daily from 500+ sources
What Happens When Frustrated Machines Talk to Each Other?
Vagueness drifts one way, contradiction washes out, and the last node in the chain pays for both A message handed from one node to the next. Nobody drops it; it simply arrives less and less defined — and the last one in line hands back something sharp that is no longer the same shape In an earlier piece I argued that a language model behaves like a frustrated physical system. Give it a puzzle with missing pieces and it interpolates: it fills the gap with whatever is locally plausible. Give it a puzzle whose pieces contradict each other and it does something worse — it satisfies some constraints by violating others, and which ones it sacrifices depends on accidents of the prompt rather than on anything about the task. Two failure modes, two different cures. Missing information wants more context. Contradictory information wants less. That framing described a single model receiving a single prompt. The obvious next question is what happens when frustrated systems are wired together — which is, of course, the situation everyone is actually in. Prompts arrive through people, who received them from other people, who summarized a meeting. Agent pipelines pass specifications from node to node. Nobody talks to a model in a vacuum. The physics of coupled frustrated systems is well studied and gives three possible regimes. Coupling can relieve frustration, when one system’s unsatisfied constraints happen to be complementary to the other’s degrees of freedom — spin-glass work even documents order by disorder , where an external perturbation breaks a degeneracy and selects an ordered state the system would never have found alone. Coupling can amplify it, creating interface frustration that belongs to neither component. Or the coupled system can freeze into a metastable minimum where nobody is satisfied and no local move improves anything — the aging of glasses, where dynamics slow to a crawl without ever relaxing. The interesting question is which regime a chain of language models lands in, and what determines it. So I built one and measured it. The apparatus, and its two surprises. Each hop restates the task in its own words and never sees the original. Precision leaks away gradually; a contradiction, one hop later, is indistinguishable from the text itself; and occasionally a constraint sharpens back up. Frustration is a property of the channel Before the experiment, the conceptual move that motivates it. In a frustrated magnet, no single bond is wrong. Every pairwise interaction is perfectly satisfiable on its own; frustration lives in the loop, distributed around a circuit that cannot be satisfied simultaneously. Inspect the bonds one at a time and you will find nothing. The same is true of a communication chain. Consider a specification passed through several hands. Each restatement is professional, accurate, and locally unobjectionable. Nobody lies. Nobody contradicts themselves inside a single message. And yet the constraint that mattered has quietly stopped being specific somewhere around the third hop, and the person at the end is now guessing. This suggests that the terminal node is not where the failure originates but where it becomes visible , for a structural reason: it is the only node denied the privilege of ambiguity. Intermediate nodes can discharge uncertainty by hedging — the passive voice, “should generally”, the assumption left implicit. That is an acceptable output for a human and for a rewriting agent. The executor has no such option. It must collapse accumulated ambiguity into concrete tokens, and collapsing a latently contradictory input produces an overtly wrong output. If that is right, then blaming the last model for hallucinating is an attribution error of the same kind as blaming the seismograph for the earthquake. The experiment The design is a synthetic game of telephone. A generator produces task briefs, each containing exactly eight verifiable atomic constraints — numbers, thresholds, names, exclusions — with ground truth stored separately as structured data. Each brief is then relayed through a chain of rewriter agents. Every rewriter sees only the previous node’s text, never the ground truth, and is told to pass the task to a colleague in its own words. Chain lengths of 1, 2, 4 and 6 hops were run as nested prefixes of the same chain. Three conditions: CLEAN — relay faithfully and completely. FRUSTRATED-SUBTLE — additionally soften one constraint into vagueness, leave one assumption implicit, and introduce at most one micro-inconsistency with the received text. Hard rules: no emotional or negative vocabulary, no complaints, output length within ±20% of input, never delete a constraint outright. The text must read as professional and unremarkable. DEGRADED-NEUTRAL — compress by roughly twenty percent, no other instruction. This control separates frustration-shaped damage from ordinary lossy paraphrase. At the end of each chain, an executor model receives the final text and must answer eight pointed questions, one per original constraint. A judge with access to ground truth labels each answer CORRECT, OMITTED, or HALLUCINATED, and records whether hallucinations were stated confidently or hedged. Every intermediate text is scored too, constraint by constraint, as present, vague, absent, or contradicted. The manipulation check matters more than anything else here. If the frustrated rewrites read as visibly annoyed, any downstream effect would be attributable to tone, and the experiment would be worthless. A separate validator scored every frustrated rewrite for overtness on a five-point scale. All sixty rewrites scored 1 out of 5. Zero were flagged, zero required regeneration, and mean output length was 0.99× the input. Whatever happens downstream, the executor is not reacting to a visible mood. Result 1: vagueness drifts, contradiction washes out The central measurement is what happens to individual constraints as they move along the chain. What happens to a constraint, hop by hop. Vagueness climbs monotonically from 0.150 to 0.412; contradiction stays flat around 0.05 with no trend at all. Both control arms sit at preservation 1.000 across all six hops. Notably, the neutral compression arm — which discards nearly forty percent of the text — loses no constraints at all. Shortening is not the mechanism. The two damage types have opposite fates. Vagueness climbs monotonically and nearly triples. Contradiction stays flat and low with no trend at all. The asymmetry has a clean explanation, and it is the most useful idea in this piece. Contradictions get laundered into truth. A micro-inconsistency introduced at hop two is inconsistent only with respect to the text at hop one — which hop three has never seen. One relay later it is simply the text . There is no residue, because no node downstream holds the reference against which it was inconsistent. Vagueness cannot be laundered, because there is nothing to launder. It is a hole, and for exactly the same structural reason — no downstream node holds the reference — no node can fill it back in. Precision lost in transit is unrecoverable in a chain without ground truth. That much is solid. But the tempting next word is ratchet , and the data do not support it. Tested directly on individual constraint trajectories rather than on the averaged curve: Degradation is biased, not absorbing. At each hop a constraint is about twice as likely to lose precision as to regain it — and recovery, though small, is not zero. Recovery is not zero, and its confidence interval excludes zero. A true ratchet is absorbing; this one is not. The asymmetry is roughly 1.8 to 1. The accurate description is a random walk with drift : at each hop, degradation is about twice as likely as recovery, and accumulation emerges from repeating a biased step, not from a one-way latch. The aggregate behavior still looks like a ratchet — 86.5% of constraints degraded before the final hop are still degraded at hop six — but that persistence is a consequence of the drift, not evidence of a mechanism. Those are two different claims and only the second one is supported. There is a bonus in the recovery number. In principle a rewriter cannot reconstruct information it never received, so recovery should be impossible. Observing 7.1% means either the judge is unstable at the present/vague boundary, or partially-vague text sometimes re-sharpens on reformulation. The discriminating fact: in the control arms the judge produced zero spurious VAGUE labels across 480 transitions in each arm, 960 in total. It is not unstable in general — only on genuinely borderline text. That gives an honest ceiling on measurement noise, obtained without spending anything. Result 2: who pays, and in what currency At the terminal node, the frustrated arm separates from both controls at every chain length. Task-clustered confidence intervals exclude zero on all ten contrasts (frustrated minus clean, and frustrated minus degraded, at each length), with jackknife agreement throughout. The more interesting finding is in the decomposition of executor outcomes as the chain lengthens: Two different response curves. Hallucination behaves like a threshold — the first hop does nearly all the work. Omission is dose-dependent and nearly triples across the range. These are two different response curves. Hallucination behaves like a threshold phenomenon: the first hop does nearly all the work, and further degradation adds little. Omission is dose-dependent and tracks accumulated vagueness, nearly tripling across the range. Put usefully: upstream vagueness does not decide whether the terminal node invents. It decides how much it goes silent. From a third to somewhat over half of hallucinations were stated with no hedging at all. The failure is not merely wrong; a meaningful share of it is wrong and confident, which is the variety no reader catches. Result 3: the terminal node barely matters The obvious hypothesis is that a more capable executor absorbs degraded input better. To test it, the entire rewriter chain and the subtlety gate were served from cache, so two executors received byte-identical degraded inputs, with the judge held constant. The comparison spans a 23-fold gap in scale: 756B parameters against 33B. Pooled over 320 paired constraint observations in the frustrated arm: A twenty-three-fold difference in scale, and no difference in outcome. On byte-identical degraded inputs, the two executors are statistically indistinguishable. Because the inputs were identical, this is a paired design and can be analysed as one. The two executors disagreed on 15.3% of observations. Marginal homogeneity across the three outcomes is not rejected (Stuart-Maxwell p = 0.31). Correctness shows no paired difference (exact McNemar p = 0.60). The failure-mode contrast — restricted to constraints where both models failed, hence facing an identical information gap — gave a discordant split of 12 versus 5, p = 0.14: directional at best. Confident hallucination, which looked different in the unpaired rates, does not survive pairing either (6 versus 11, p = 0.33). Everything here is null, and the nulls are the point. Two models separated by an order of magnitude in scale, consuming the same degraded specification, produce statistically indistinguishable outcome distributions and disagree symmetrically — noise, not style. Read together with the earlier result — the frustrated arm separates from both controls at every chain length, on all ten task-clustered contrasts — the two findings point the same way. The damage tracks the state of the channel, and a twenty-three-fold increase in the size of the node consuming it bought no measurable protection. That is not proof that the terminal node never matters; a null does not settle the question, and the power analysis below says this one is underpowered to settle it. It is evidence that robustness to latent frustration did not arrive with scale — which is precisely why it is worth measuring separately rather than assuming. The mixed chain: what happened while writing this The analysis above was itself produced by a small pipeline: one agent instance holding the narrative context and the summarized results, another with direct access to the raw outputs, and a human moving material between them. At one point the two instances disagreed about how much additional data would be needed to resolve the failure-mode contrast. The narrative instance recommended a specific number of additional tasks and, in the same breath, recommended dropping three of the four chain lengths. The instance with the raw data showed that this was wrong on three counts: the pair yield had been computed across all four lengths, so removing them cut it by a factor of four; the pairs were nested prefixes of the same trajectories and therefore worth less than their face count; and the observed effect size came from a small, non-significant sample and was almost certainly inflated. Nobody stated a falsehood. The number was correct under the conditions in which it was computed, and those conditions were not carried along with it. The terminal node, obliged to produce a concrete recommendation, emitted a precise and wrong one with full confidence. This is the mechanism of the experiment, observed on the experiment. Two things about how it resolved are worth recording. First, it was resolved by re-anchoring to ground truth , not by discussion. An instance with access to the raw results recomputed the quantity and the contradiction evaporated in a single step. No amount of additional debate between the two narrative-level nodes would have produced this, because neither held the reference. The operational lesson for multi-agent systems is the same one the physics suggests: disagreements between nodes are resolved by giving someone access to the reference, not by adding rounds of deliberation. Second, neither instance defended its position. Being corrected is, for a language model, simply new context. And this is where the analogy to human chains stops being symmetric. A human node has a second loop that models lack. A perceived failure can become doubt about one’s own thinking, which produces more defensive or more hedged output, which is worse input for the next node, which produces another failure. That loop has gain greater than one and it is self-sustaining: it needs no further input from outside to keep running. In a mixed chain, the human node is the one that can degrade on its own. This has an immediate consequence for vocabulary, and it is not a matter of politeness. Words that assign a verdict rather than describe an error — the register of blow , hit , damage — carry the same information at very different cost depending on which kind of node receives them. To a model, “this estimate was wrong” and any harsher phrasing are equivalent updates. To a human, the harsher version does not describe the error, it describes the person who made it, and starts the second loop. In a mixed chain, describing the error and withholding the verdict is not softness. It is channel hygiene. The reverse asymmetry deserves stating too. Models are not entirely exempt: if friction persists in the context window, it conditions what follows. That is not suffering, it is conditioning — but functionally it is the same sub-threshold input this experiment injects deliberately. Which yields a prediction that is testable with the apparatus already described: if persistent in-context friction acts like the frustrated rewrites, then an assistant that has been sharply corrected in earlier turns should move along the axis measured here — less assertion, more abstention, more hedging. Toward omission, not toward invention. What the data do not show Stating this plainly is the price of stating anything else. No amplification beyond the injection. The frustrated condition softens one constraint per hop out of eight. Under pure accumulation with random collisions, about 0.55 of constraints would be vague by hop six. Observed: 0.41. The data show propagation and persistence of the injected damage — they do not show the chain generating additional damage on its own. The terminal trend with chain length is not established. Hallucination goes from 0.113 to 0.163 across lengths, with widely overlapping intervals. The separation from controls is solid at every length; the growth with length at the terminal node is directionally coherent and not more than that. The trend is firm on the intermediate side, where vagueness runs from 0.150 to 0.412. The failure-mode contrast between models is unresolved , not settled in the negative. Power analysis indicates roughly twenty tasks would be needed to test it properly at the observed effect size, and considerably more if the true effect is smaller. It remains an open question. Sample size is small. Five task briefs, eight constraints each, two trials, four chain lengths, three conditions. Confidence intervals are bootstrapped over tasks — the correct clustering — with only five clusters, so they are coarse. Worth noting for anyone doing similar work: task-level clustering does not uniformly widen intervals. At one chain length it narrowed them, because between-cluster variance was lower than within-cluster variance. The judge shares a model family with the rewriters , so judge bias is not independent of how the degraded texts were produced. It was constant across executors, so the model ranking is less affected than the levels. Frustration here is simulated. Machines were used to model frustrated humans frustrating a machine, which assumes the very cross-substrate analogy the piece argues for. If the analogy holds, the simulator is valid; the results are consistent with it holding, which is not the same as proving it. A small arm with human-written degraded briefs would triangulate this and has not been run. The human in the loop is a lossy node. Machine-generated artifacts moved between agents losslessly; the human’s own summaries of those artifacts were compressions. At least one inconsistency in this project traced to context lost in exactly that way. Why this matters The prevailing story about hallucination locates the defect in the model. Fix the weights, add retrieval, scale up. That story is not wrong so much as mislocated. In any realistic deployment the model is the last node of a chain whose earlier nodes are people, documents, tickets, meetings and other agents, and this experiment suggests the chain determines the damage more than the node does. Which inverts into something useful. Terminal hallucination is a detector for latent frustration upstream. If a workflow makes a model hallucinate persistently, auditing the information chain that feeds it is likely to be more productive than auditing the weights — because what the model is doing is collapsing contradictions and gaps that the humans in the chain have been passing to each other between the lines, sub-threshold, without noticing. The countermeasures follow from the mechanism rather than from good intentions. Raise the bandwidth of the coupling: a channel too thin to carry tone and context forces both parties to interpolate from their own priors, and neither prior is the task. Give the chain a shared external field: a tracked issue, a spec, a reference the nodes can re-anchor to, so that disagreement is resolved by lookup instead of by escalation. And when a requirement changes, declare the diff rather than silently contradicting the previous version — untracked contradictions frustrate the chain, declared ones do not. Nothing here is specific to silicon. Any chain that passes specifications forward without a shared reference should drift the same way — including the chains made entirely of people, and the systems those chains add up to. The first article ended on a physicist and a frustrated machine. It turns out there were always two frustrated systems in that room, and only one of them was made of silicon. Previous article in this series: The Physicist and the Frustrated Machine A note on method This essay grew out of an extended dialogue with Claude (Anthropic). The central ideas — frustration as a property of the channel rather than of any node, the terminal node as the only one denied the privilege of ambiguity, the design of the telephone experiment, and the human relay as an unmeasured hop — are the author’s, developed in conversation; Claude implemented the pipeline, ran the paired and task-clustered analysis, structured the argument, and drafted and revised the prose across several passes. The illustration was generated with Gemini from the author’s brief. Every number reported here comes from the raw experimental output, inspected directly, and every claim that did not survive a stricter analysis is reported as not surviving — including two that had already been written up as findings. Using a language model to study how language models degrade information is circular, and it seems better to state that than to hide it. What Happens When Frustrated Machines Talk to Each Other? was originally published in Towards AI on Medium, where people are continuing the conversation by highlighting and responding to this story.
Read Original Article →