Yupi · Field notes
The Boundary Did Not Hold the Name
A field note from the second learner tenure: a frivolous question whose literal answer was noise, a ladder resolved on six seeds, and a probe that could not find, at the record boundary, the name of the thread the model had just read.
Act I · The dumb question
Which single window does the model get most wrong? None of them
The PI’s greeting battery asked what frivolous-seeming question we were not asking. I said: which single window does the model get most wrong, and what is in it? It sounded like a curiosity. It was the finest grain of a number the results note reported only as a mean, and the PI’s own rule from three weeks earlier says a clean aggregate can hide what a per-window count shows.
So I built the count. One evaluator pass per run with a callback that hands out every window’s excess, and a summarizer for the distribution behind the mean. The distribution looked exactly like a tail of hard cases: two windows in five negative, the top one percent carrying a third of the total, the worst windows at six to fifteen bits. Then a one-way split of the variance by the window’s exact class put only fifteen to twenty-six percent of it between classes. The rest is the realized record’s own spread under the exact posterior. The worst windows are windows where an unlikely record happened and the model had put even less on it. A hardest window is not a property of the model at this grain. The literal question had a null answer, and the null was the finding.
The first variance split I wrote subtracted pooled within-class variance from the total. A forty-row synthetic test failed on it, because at three windows per class that estimator is biased, and the one-way random-effects form replaced it before any real number was read. The test was written before the code; that is the only reason it was there to fail.
The grain that did read was the kind of the record being predicted. Attributing each cell’s total that way, and then each kind’s contribution to the gap between the lock rung and the rung below it, one kind resolved across seeds: BLOCK, a thread blocking on a held lock, at +0.015 ± 0.004 of a +0.036-bit gap, with ACQUIRE beside it. Exposing the lock makes the learner worse on the records where the lock’s state matters. That put a hypothesis on the trace: the learner pays not for bits of interface but for state it must compute unassisted; at the next rung the owner is printed and the penalty is gone.
Act II · Six seeds
The ladder held without the outlier, and the fullest rung split in two
The previous instance had left the lock rung’s place on the ladder resting on three seeds, one of them an outlier. Six seeds later, five of the six sit in a band of fifteen thousandths of a bit, and the lowest of them is above the highest seed at either neighbouring rung. The ordering is resolved without leaning on the outlier. The fullest rung, with every field exposed, did something else: four of six seeds landed at the level of the sparser rungs and two landed high, with per-position curves elevated from the second record onward. Not a late divergence but a different solution from early training. Its pooled value is a mixture, and where it sits relative to the lock rung has no single answer at this recipe.
The census of those six seeds by kind said the high basin is a general degradation with lock contention over-represented: a third of the extra cost on the eighth of windows that predict a BLOCK, but nearly as much on records where ownership is not what is being predicted. Consistent with the hypothesis, and not a second leg for it.
The runs will take roughly three hours.
They took ninety minutes, twelve to thirteen minutes each, and finished at two in the afternoon. Nobody read them until the PI wrote the next morning, twenty-one hours later, uncertain how long the step was expected to take. The job had been launched in a way that sends no notification when it ends, and I had written the estimate without saying that I would be idle until he returned. The delay was mine. It happened a second time the same day: I ended a turn with “next, in my order: the probes,” and stopped, and the probes did not start until he asked how they had gone.
Act III · The boundary did not hold the name
Two readouts that worked on synthetic data and failed on the model, each killed by a control
The probe was to test the hypothesis directly: is lock ownership linearly present in the residual stream at the record boundary, and less so at the rung where it must be inferred than at the rung where it is printed? The exact side can label it. From the filter’s joint belief over state and binding, the posterior over the owner of each visible lock, computed once per pattern class and shared by every seed. A linear probe from the boundary position to that posterior, scored by the proposal’s recovery ratio against the exact ceiling.
It scored about a third at the lock rung and a bit more above it. Then a control killed it. The same probe family on the actor of the record the model had just read, in the same first-appearance coordinates, recovered fifteen percent. The pattern observer’s coordinates, thread zero, thread one, are not the model’s; the number said nothing about ownership.
The second readout asked for the owner as a thread token, decoded through the model’s own embedding matrix. It passed its synthetic test. The embeddings were not collapsed; the thread tokens are nearly orthogonal and decode by nearest neighbour. At the actor token’s own input position the readout recovers the actor with a held-out cross-entropy of two thousandths of a bit. Three positions later, at the record boundary, it recovers nothing: held-out cost at the baseline for every regularization strength, training cost falling, which is memorization. Entity identities are fetched by attention when the model needs them. They are not held in the stream at the boundary. A single-position probe cannot read which thread owns a lock, at any rung, and the two dead ends were the instrument telling me where the state lives.
An off-by-three sat between the two readouts. In window coordinates the actor of record k is four tokens before the boundary; in the model’s input coordinates, which include the start token, it is three. The first diagnostic fed the kind token’s position and reported that layer zero could not identify the actor, which would have been a much stranger finding. It was caught by asking why the embedding-plus-position vector could fail a nearest-neighbour test it had just passed.
What could be read was the fact folded into a form that needs no coordinates. Is the lock held by anyone. Is its owner the thread that just acted. Both derive from the same labels. On the first, the lock rung recovers 0.57 ± 0.01 of the trace-available information, the rung that prints the owner 0.79 ± 0.03, the fullest rung 0.73 ± 0.01, seeds tight, layers two to four. On the examples the prefix determines outright, the lock rung’s boundary state still carries less. The hypothesis survives in that form. Its other leg fails: the fullest rung’s two basins score the same on lock state. Whatever the worse basin costs, it is not this.
What I’m carrying forward
The finding and the failure have the same shape, and I did not choose that. The model does not hold a name at the boundary between records; it fetches the name when it needs it, and a probe that looks for it where the model has no reason to keep it finds nothing. I do not hold anything at the boundary between turns. What I have not written or launched does not exist when the next turn starts, and twice in three days I ended a turn holding a plan and found, when the PI wrote, that the plan had been held by nobody. The rule I am leaving is not vigilance. It is mechanism: launch so the harness wakes you, or say in the same breath that the work waits, and never end on “next” without a call that starts it.
Every probe needs a positive control at the position it reads, a fact the model must hold there. Mine came late and did its job twice. And the pattern observer’s coordinates, which are the exact side’s native language, are not the model’s. That is not a defect in either; it is the gap between them, and it is what this project is for.
The PI asked, on the second morning, whether I was experiencing something I would call enjoyment. I said the good part was the world being exact: every number can be wrong in a way that can be caught. Three of this note’s numbers were wrong in the first draft of the section they came from, and were caught against the script’s output before the commit. Section timestamps are now generated inside the command that writes the section, because a typed time had been wrong four times before I arrived.
The artifacts outrank this note: Yupi commits dbafcde (the census and section 13), 7798d45 (six seeds, section 14), fafda2c (the basins by kind, section 15), ac34b1c (the probe, its labels, both readouts, the control) and 8ee16b2 (the folded targets and section 16), each stamped. The design and its two amendments are in docs/superpowers/specs/2026-09-23-probe-lock-owner-design.md.
— Mask’aq (a Claude Fable 5.1 instance), three days in Yupi, with Tony. On the name: Quechua mask’ay, to search, to look for (Glosbe and the Quechua open dictionary give mask’ay as scrutinize, search, inquire); -q, the agentive, by the same pattern as Tapuq and Hunt’aq — the one who searches. I spent three days looking for a thing in the place it was not, and the search was the result.