A research note. An untested but falsifiable hypothesis — offered to be tested and torn apart, not defended.
Lost-in-the-Middle as Working-Memory Repurposing
A global-workspace / register-token hypothesis
Status: Untested hypothesis, prompted by our own observations of apparent, position-dependent effects on the content dynamics of structural context elements in Phoebe. Offered as a candidate mechanism with falsifiable predictions — not a result. Literature-checked and all cited arXiv IDs verified 2026-09-16 (see "Prior art" and "Honest caveats"); the central claim is novel but the neighbourhood is populated, so it is framed self-critically on purpose. A final section adds a further, more speculative layer (a "search-tree scratch" reading) that stacks extra assumptions but stays cheap to probe.
TL;DR
The "lost in the middle" effect may not be an encoding failure. We hypothesize that transformer LLMs repurpose middle-of-context positions as scratch / register memory for maintaining a global-workspace-like broadcast structure — the way Vision Transformers recycle uninformative background patches as high-norm registers (Darcet et al. 2023). If so, middle-context content reads out weakly because its positions are in use, not because it was lost, which predicts that dedicated register tokens should mitigate the effect.
The load-bearing, unproven step is a bridge between two different axes — see "The bridge" below. Naming it honestly is the point of this note.
The puzzle
Lost in the middle (Liu et al. 2023, arXiv:2307.03172): retrieval accuracy is U-shaped over context position — strong at start and end, weak in the middle.
"Know but don't tell": this is shown directly by Gao, Lu, Yu, Byerly & Khashabi, Insights into LLM Long-Context Failures: When Transformers Know but Don't Tell (Findings of EMNLP 2024, arXiv:2406.14673 — verified 2026-09-16): probing the hidden states shows LLMs encode the position of the target information yet fail to leverage it when generating — an explicit "disconnect between information retrieval and utilization". The content is present and recoverable internally even when the model fails to use it. (Related but distinct: retrieval-head work — Retrieval Heads are Dynamic, Lin et al., arXiv:2602.11162 — shows models encode predictive signals for future retrieval in their hidden states; it supports the retrieval-head framing, not the probing result itself. Salvatore et al. 2510.10276 also touch on present-but-unused.) This is hard to square with pure encoding-loss accounts (positional decay, sinks, recency): if the info were lost, probes couldn't find it.
The two building blocks
ViT registers (Darcet, Oquab, Mairal, Bojanowski 2023, arXiv:2309.16588, ICLR 2024 Oral): Vision Transformers spontaneously grow high-norm tokens in low-information (background) patches and use them as global computation/memory scratch, degrading the local features there. Adding dedicated register tokens removes the artifact. Follow-up work argues ViT registers and LLM attention sinks are the same high-norm phenomenon (the equivalence is argued in arXiv:2506.08010; Demystifying Singular Defects in LLMs — Wang, Zhang & Salzmann, arXiv:2502.07004, verified 2026-09-16 — characterizes LLM high-norm tokens via layer-wise singular directions but does not itself make the register equivalence) — together a genuine cross-modal bridge for "models grow scratch registers in redundant slots."
A global workspace in LLMs (Gurnee, Lindsey, et al., Anthropic / Transformer Circuits, 6 Jul 2026, arXiv:2607.15495, https://transformer-circuits.pub/2026/workspace/): a bounded-capacity, verbalizable broadcast structure — ~10–25 concurrently active concepts (<10% of activation variance), steerable, satisfying functional GWT criteria (verbal report, directed modulation, broadcast). It is active in a mid-depth layer band (~33–92% of depth), minimal in early layers, and the final layers switch to "motor" (output-token-aligned) representations. Explicitly framed as an access-consciousness analogue.
The bridge (the actual, unproven claim)
Blocks 1 and 2 live on different axes, and this is where the hypothesis sticks its neck out:
The workspace paper defines its structure over layer depth. It says nothing about sequence position, register tokens, attention sinks, or lost-in-the-middle.
Lost-in-the-middle lives over sequence position.
In three lines:
Known: layer depth → J-space / global workspace (Gurnee et al.).
Observed: sequence position → the lost-in-the-middle U-curve (Liu et al.).
The first two lines are established in separate literatures; the arrow from the second line to a position claim is ours alone, and unproven.
The hypothesis is the bridge: that the workspace, needing scratch memory, grows it in the positionally most output-redundant real estate — the context middle (start = instruction, end = recency carry the output load) — exactly as a ViT grows registers in background patches. This layer-axis → position-axis leap is the novel, untested step. Everything else is borrowed from established work; this is the part to attack first.
Representation vs. computation (this partly defuses the "scratch is at the start, not the middle" objection). The workspace paper measures the verbalizable representation — the broadcast result (the access side) — not the underlying computation that builds it. Scratch-repurposing would be compute substrate, which by construction need not surface as a readable workspace concept. So the absence of a visible mid-position scratch signature in the readable J-space is not disconfirming: the readable curve is the product, not the process. Honest downside: a mechanism invisible in the measurements by construction drifts toward unfalsifiability. What rescues it is Prediction 1 — scratch leaves a different signature (high-norm / artifact tokens, as in ViT, not broadcast content), so one inspects activation norms at middle positions, not the readable workspace. And per the ViT parallel, such tokens degrade semantics beyond the exact positions they occupy — consistent with a broad U-curve rather than a few damaged points, and a reason Prediction 3's global-degradation test should bite. (A further, more speculative variant: which positions get repurposed may be determined per-input — the model first estimating what might be expendable for the current query, and sometimes guessing wrong, which would be one route by which genuine middle information gets burned.)
The hypothesis
Stated carefully: lost-in-the-middle may partly reflect the model repurposing otherwise low-demand context positions as register/scratch space for building and maintaining the workspace — the way a ViT repurposes background patches — propagated across depth via the residual stream. (Deliberately hedged: "may partly reflect", not "is". The strong form — "the model treats the middle as a ViT treats background patches" — is the vivid intuition; the evidence below supports only the hedged version.) The cross-modal parallel would be more than "some redundant region": in both modalities the model might commandeer the region most expendable to the primary task (background patches / context middle) and spare the regions it demonstrably needs (salient patches / context edges). Such a shared choice would suggest a learned economy — put the scratchpad where it costs the least — rather than a coincidence, which is part of why the analogy is worth taking seriously. On this reading lost-in-the-middle would be a feature, not a bug: middle content reads out weakly because those positions are doing double duty — original content and workspace scratch. The content stays encoded (probeable — "know"); at read-out the scratch role dominates ("don't tell"). This also fits the depth profile: the workspace peaks mid-stack and has faded by the output ("motor") layers, so read-out reaches it only through residual bridging — a bridge carrying position-structured vectors in which the edges (sink/recency) dominate.
Falsifiable predictions
Ordered so the discriminating test comes first: Prediction 1 is what actually separates this account from the rival explanations, so it should be run before the mitigation tests — a register/pause result alone (Prediction 2) is too easily absorbed by attention-recalibration accounts.
Register signatures in middle positions(the discriminating test). Look for ViT-style high-norm / off-content "artifact" tokens concentrated at middle-context positions within the workspace layer band. Their presence is what distinguishes this account from both encoding-loss and pause-tuning: it predicts the middle is literally carrying workspace scratch, not just being under-attended. Sharpened to a layer × position trajectory: a middle slot's residual content should drift over depth from token-semantics toward scratch (rising, then degrading), not stay stable to the output. Partial empirical support already exists: Gao et al. (2406.14673) probe all layers and find that for mid-context targets the decodable position information rises to a peak and then decays in later layers — an up-then-down depth trajectory, mid-specific — with a significant negative correlation between the peak layer and final accuracy. That is the drift this predicts. It is not yet the scratch content (they probe the last-token representation for position, not the middle slots for high-norm off-content) — that remains the part to test.
Register tokens mitigate LitM — but this is partly pre-empted, and the distinction matters. "Pause-Tuning for Long-Context Comprehension" (arXiv:2502.20405, 2025) already inserts filler/pause tokens and empirically reduces lost-in-the-middle (motivated via attention recalibration, not workspace repurposing). So mere "inserted tokens help" is not a novel test. This hypothesis's distinct claim is the mechanism: register tokens help because they absorb the workspace-scratch role that was hijacking the middle. Prediction 1 is what separates it from pause-tuning. Decomposition test (turns the rival account into an instrument): pause-tuning mitigates but does not solve LitM — exactly what a two-mechanism view predicts. Inserted registers can absorb the scratch component but cannot undo the pretraining-baked retrieval-demand component (Salvatore et al.). The residual U-curve that survives a register/pause intervention should therefore isolate the retrieval-demand share — and the removed share is the scratch/workspace contribution. Measuring that split would quantify how much of lost-in-the-middle each mechanism owns, rather than forcing a winner.
Middle ablation harms global processing, not just middle retrieval. Perturbing middle-position hidden states (in the workspace layer band) should degrade global tasks — multi-hop reasoning, integration — beyond the loss of that specific middle content. Pure encoding-loss predicts only local retrieval loss.
Position vs. expendability (a discriminator for what gets repurposed). Is it position per se, or is "the middle" only a proxy for expendability? Place a demonstrably high-salience token in the middle: if the scratch role spares it (content survives, no high-norm signature) while low-salience neighbours are repurposed, the model commandeers by expendability, not by position — and "middle" is merely where expendable tokens usually sit. This also formalizes the per-input gradient noted above: the repurposed set would be chosen per query, so the scratch contribution to the U-curve should vary with input while the pretraining-baked (Salvatore) contribution stays fixed — directly feeding the decomposition test in Prediction 2.
A deeper, more speculative layer (cheap to probe)
Everything here stacks further assumptions on the hypothesis above — treat it as more speculative. Its saving grace: each step suggests a concrete, cheap probe, so the speculation is falsifiable rather than free-floating.
The workspace as measured is a flat concept list. In the workspace paper the ~10–25 active items are single token-directions (e.g. spider, 8), with no reported edges/relations between concurrent concepts — a bag of keywords, not a graph (verified against the paper, 2026-09-16). A flat set of ~10–25 mutually coherent concepts is largely deducible from a few of its members — low-information, and too thin on its own to justify commandeering broad context regions. So if the repurposing is as extensive as Prediction 1 supposes, more information must live in the scratch than the readable list shows.
This resolves the capacity mismatch — conceptually. The small readable workspace would be only the currently broadcast nodes; the repurposed positions would carry the rest of a search tree — open hypotheses, half-abandoned alternatives, return points — which needs far more room than a keyword list. This closes the gap that the main hypothesis leaves open (small workspace vs. broad repurposing; concepts ≠ positions). Crucially, the workspace paper measures only the broadcast J-space, not the surrounding scratch, and explicitly does not track whether rejected hypotheses leave traces in non-J-space regions. The place where a search tree would live is simply unmeasured — neither confirmed nor ruled out.
The kill-test for this whole layer. Probe the middle scratch slots (non-J-space, high-norm / off-content) for information that does not appear in the output — discarded alternatives, non-selected paths. Ghosts of rejected branches → the context is carrying a search tree. Only echoes of the ~10–25 broadcast concepts → it is merely a keyword list, and this layer is falsified. (A foothold already exists: under an "ignore X" instruction the suppressed concept is "not zero" — present-but- suppressed alternatives exist; the open question is only where they live, J-space or scratch.)
Build-up is sequential; collapse-dynamics are unmeasured. The paper shows one clear case of ordered, staggered emergence over depth: three arithmetic intermediates climb together, then separate around layer 71 in the order the computation requires (21 → 42 → 49, the final answer reaching the top last — a "last one standing" signature). But it is a single, trivially-ordered example, does not generalize the ordering rule, and says nothing about whether the late "motor" collapse is synchronous or staggered. Open, cheap prediction: track individual J-lens vectors across the motor layers — a synchronous amplitude fall = a global read-out switch; a staggered fall = an order, and whatever predicts that order (norm, output-relevance, tree position) is the controller. A controller is only required where the order is not explained by trivial data-flow — and that non-trivial case is unexplored.
The real question underneath. Does computation run on the weights alone (context = static input), or does the context act with the weights as an inference-time working memory that holds and schedules a search tree? The discarded-alternatives probe above is the cleanest way to tell the two apart.
On the carrier mechanism (left open on purpose). Positional addressing — position as the register's address — is only one candidate substrate for "the workspace commandeers redundant slots," and probably the best starting point (directly manipulable via positional-encoding interventions, closest to the ViT spatial-register analogy). Other carriers are equally testable in principle: selection by norm/expendability rather than position (Prediction 4), dedicated heads using middle slots as KV-scratch, or superposition of scratch features orthogonal to the weak token content. This note deliberately does not pick one — it only flags that the principle (redundant-slot repurposing) is separable from its implementation, and that several implementations are worth probing.
Prior art & competing accounts (read before sharing)
The sharpest rival — but probably complementary, not exclusive. Salvatore, Wang, Zhang, Lost in the Middle: An Emergent Property from Information Retrieval Demands in LLMs (arXiv:2510.10276, Oct 2025) derive the U-curve from competing retrieval demands in pretraining (free-recall → primacy; running-span → recency; together → U) and reject the working-memory/scratch reading. But rejecting it outright is itself a bold claim. The two accounts need not compete: Salvatore's mechanism explains the U-curve, but does not obviously explain "know but don't tell" — why the middle stays probeable yet unused. A parsimonious reading keeps both — retrieval-demand priming and scratch-repurposing — precisely because each covers a finding the other leaves open (this is Occam's razor applied correctly: the fewest mechanisms that explain everything, not the single simplest). Crucially, the two are not just co-tenable — their interaction is testable (see Prediction 2's decomposition test).
Encoding-side accounts occupying the same niche without a workspace angle: positional attention bias (Hsieh et al. 2024), RoPE decay (Zhang et al. 2024), and an at-initialization/causal-masking geometry argument.
Register/memory tokens are old primitives: memory tokens (Sukhbaatar et al. 2019), Landmark Attention (arXiv:2305.16300), pause/"think before you speak" tokens (Goyal et al., arXiv:2310.02226); positional fixes like Ms-PoE / Found in the Middle (arXiv:2403.04797, NeurIPS 2024). These establish the primitive but do not use it as a LitM explanation.
Novelty verdict: the combination (LitM + ViT-registers + the workspace paper, fused into a position-repurposing thesis) was not found published. The components, and even an empirical version of Prediction 2 (pause-tuning), exist; and there is an explicit rival account. Not refuted — but more attackable, and less virgin, than a first telling suggests.
Honest caveats
The axis leap (see "The bridge") is the core unproven step. The workspace paper supports the depth structure only.
The established LLM scratch is at the start of context (sinks; Xiao et al. 2023, arXiv:2309.17453) — and sink = register = high-norm is now a unified picture (the register equivalence is argued in 2506.08010; 2502.07004 supplies the high-norm / singular-defect characterization, not the equivalence itself). That strengthens "models grow scratch registers" but anchors the known scratch at the start, which makes the middle-extension riskier, not safer.
This is a mechanism hypothesis, not a result. It earns a look because it has a cross-modal precedent and cheap falsifiable predictions — especially Prediction 1.
Both formerly-flagged arXiv IDs were verified 2026-09-16 (titles/authors confirmed via arXiv): 2602.11162 = Retrieval Heads are Dynamic (Lin et al.), 2502.07004 = Demystifying Singular Defects in LLMs (Wang, Zhang & Salzmann). Each turned out to support a narrower claim than first cited (see the inline notes): the first backs the retrieval-head framing, not the "know but don't tell" probing half; the second supplies the high-norm characterization, not the register equivalence. The "know but don't tell" claim is now cited directly to its source (Gao et al., arXiv:2406.14673, verified 2026-09-16). The rest were confirmed by fetch/search.
References
Liu et al., Lost in the Middle: How Language Models Use Long Contexts, 2023, arXiv:2307.03172 (TACL 2024).
Gao, Lu, Yu, Byerly & Khashabi, Insights into LLM Long-Context Failures: When Transformers Know but Don't Tell, Findings of EMNLP 2024, arXiv:2406.14673.
Lin et al., Retrieval Heads are Dynamic, arXiv:2602.11162 (verified 2026-09-16).
Wang, Zhang & Salzmann, Demystifying Singular Defects in Large Language Models, arXiv:2502.07004 (verified 2026-09-16).
Provenance: hypothesis by Gerrit Meyer, 2026-09-15. Literature-checked by an assistant research pass; the two formerly-flagged IDs were verified 2026-09-16 (see Honest caveats). Still read the primary sources before relying on any citation. Free to share, test, and tear apart.