← Oscar Labs

essay

row four was always there

31 August 2026

TLDROne night of coding agents left ten surfaces carrying a sentence their own cited source contradicts, across nine repositories, and not one was caught by the agent that wrote it. In the clearest case a merchant tool announced that no published figure priced an obstacle; the figure was row four of the same table it was already quoting two rows from. Then the piece itself said nine instances over a list of ten, and only a reader who had not written it counted the bullets.

Check it yourself
  • https://morkeeth.substack.com
  • git show 3b604ba:.gitignore
  • transcripto.py:799
  • git cat-file -t c2b1ad98112c5b9b67f888b7616c5bee18f63501
  • bin/zup.js:307
  • electron/lib/project-state.ts:85
  • shared/timeline.ts:248

In one night a run of coding agents touched nine repositories, and left ten surfaces carrying a sentence that the surface's own cited source contradicts. I opened all ten sources myself and pasted the commands below. None of the ten was caught by the agent that wrote it. Every one was found afterwards, by someone who had not written it. And the sources were never hidden: in the clearest case, the figure a merchant tool announced did not exist anywhere was row four of the same table the tool was already lifting two other numbers from.

That last part is the finding. This is not a knowledge problem, an attention problem, or a model problem. The evidence was one click away in every single case, and the person one click away from it was the only person who did not click.

The operator had asked the question himself, an hour and fifty minutes before the run started, about a completely different piece of work. Typed at 22:04 Paris time, typos his:

are we comfortable with the recommendations we're giving? are the yverified?

Then at 23:40, winding the day down:

ok i know we close down, but im gonna have my mac open, could you schedule a night run here with ambition? do we have enough to do, feedback and decisions made? how would it look, til tomorrow at 11 am, 12 hour run.

Four more turns follow that one. The last thing he typed was at 23:48:41, and then nothing, which one of the night's own builds independently drew as the handoff on its activity card. The run fanned out in waves until morning. Every wave answered his first question with a yes, in writing, in a closeout with pasted output under it, and every one was wrong about at least one sentence.

what counts as one

Before the count, the unit, because a count without a unit is the same defect one level up.

An instance is a declarative sentence, on a committed or shipped surface, that rests on a specific named source, where opening that source contradicts it, and that was present at the moment its author wrote the closeout calling the build finished.

That definition throws out a lot of real defects from the same night. A page that recursed forever, a strip that clipped five cards out of seven, a security test grading a process that was not the product: all real, all shipped, none of them this. It also collapses several sentences in one product down to one row, because two false sentences resting on the same unopened page is one unopened page. Under a looser unit the number is larger. The unit is the reason the number is ten, and it is written down before the counting so it cannot be tuned afterwards. Ten instances, on ten separate surfaces, in nine repositories: one repository carries two, an essay and a launch note that rest on different sources and were written eight hours apart, so they do not collapse into one row.

The ten also split in a way worth stating, because a headline that says "in one night" can quietly imply all ten were written that night. Four were: the account-wall sentence at 00:11, the collapse sentence at 00:24, the activity card's label at 01:54, the registry shelf at 02:12, all on 31 August, all caught within about an hour. The other six were already standing. The competition paper's reproducibility claim dates from 10 August, and the night's own work copied it verbatim into a new canonical document at 00:01 without re-reading it. The board's un-awaited probe landed on 16 August. The essay's headline dates from 27 August, the launch note from 28 August, the prompt grader's overlap was published to the package index on 28 August, and the console's records were written on 29 August. Those six are the worse half. A sentence written at 00:11 and caught at 01:21 cost seventy minutes; a sentence standing since 16 August was carried past every pass that touched that repository, and each pass certified it again by not looking.

row four

A merchant readiness tool scores an online store on how well a shopping agent can buy from it. It deducts points for each obstacle, and the whole pitch is that the deductions are not invented: each one is priced at a figure published by a cited source, and the product prints the citation on screen next to the number.

The tool shipped this, in two components, both of them on the screen a judge opens:

No published figure prices an account wall separately. It closes the same door as a CAPTCHA, so ReadyCounter charges it the same 24 points and says so.

That is an admission, not a boast. The build's own closeout even flagged the line as the weakest thing in the product: "it is defensible; it is also the sentence a hostile judge would pull on." The author looked directly at the sentence, judged its risk, and did not open the page it was a claim about.

I fetched that page. 2026-08-31, 02:23 UTC, curl, raw HTML, 80,292 bytes, HTTP 200. It carries a table headed "Causes of Agent Cart Abandonment" with six rows:

Stale price or stock data at checkout 26% Captcha or verification wall 24% Price mismatch vs listed feed 18% Required account or login 15% Unsupported payment method 11% Ambiguous page structure 6%

Row two is where the product's 24 comes from. Row one is where its 26 comes from. Row four prices an account wall separately, at 15%, which is the figure the product told its reader does not exist.

The mechanism is visible in the repository. At the commit where that sentence shipped, the research file the product cites had reproduced exactly two of the six rows: 26% and 24%. The gap the sentence reasoned from was a gap the build had made itself, four files away, by copying only the rows it already needed. Then it reasoned from its own copy as though the copy were the source.

Fixing it was not a word change. An account wall is now its own scored line at its own published 15, and a store carrying both walls pays 39 instead of 24. I ran the product's own verifier rather than take that from the fix note: a store carrying both walls pays both, not one, 94 to 55 = 39 pts. One demo store moved from 57 out of 100 to 71, and the stale 57 was printed three times in the submission draft and twice in the demo script, counted as occurrences rather than as lines, because counting lines is the same error one level down.

the other nine

Each row is a sentence a product printed, the source that contradicts it, and the command I ran. Read times are 2026-08-31 CEST unless the command says UTC.

A design page. Its one device is a collapse: fifteen task cards laid out five wide, then folded into one column, and the whole argument rests on the two drawings being the same drawing at two widths. The page said so: "The same 15 plates are in both. Only the width changes, from five to one." The stylesheet is in the same file as the sentence. git show f9cadcb:…/index.html and count the rules scoped to the collapsed state: fourteen, of which twelve apply at the width the argument is measured at, and three of those twelve touch a column track. The other nine hide elements, change type size, or change spacing. The page's own author had the contradicting source open in the same editor window.

A competition paper. Line 513: "Everything reruns from a clean checkout." git show 3b604ba:.gitignore line 5 is sdk/, and git ls-tree -r --name-only 3b604ba | grep -c '^sdk/' returns 0. The five commands the paper hands a reader fail on a clean checkout with ModuleNotFoundError. The paper's subject is numbers that are correct about the wrong object.

An activity card that draws a night of agent work the way a running app draws a run. Its honesty paragraph exists to explain why counting raw records overstates human involvement, and inside that paragraph it printed a number as "lane briefs the orchestrator wrote to its own subagents." git show 2c45d90^:agentgrinder/fleet.py:365 is sub_user_records += len(win(s["user_stamps"])): every record of that type in every lane transcript, most of which are tool output returning to the tool that called it. On a frozen window I re-derived both numbers myself: about 2,414 records would have printed under that label, and 30 of them are actually a brief. The build's own control, run separately, got 2,182 and 22. Two readings that disagree about the number and agree completely about the finding.

An essay, the one directly before this in the same series. It opened: "prompts that state what done looks like reached a durable commit 65% of the time, 64 of 99." The instrument it cites defines the word three times on the same page, including in a block of its own pasted output: survival means a durable write or an un-reverted commit. transcripto.py:799 is _DURABLE = ("commit", "artifact"). Re-ranked on commits only, that row is 32 of 108, 29.6%, against 63.9% under the definition the piece actually printed. I re-ran the script myself rather than take the receipt: 2,917 files, 414,213 records, 2,058 episodes, same table.

A claim-witness console,** whose product thesis is that a probe aimed at the wrong object manufactures a false result. On the live service right now, two of seven records carry an identical 40-character hex value in both the session field and the head_sha field. git cat-file -t c2b1ad98112c5b9b67f888b7616c5bee18f63501 returns commit, dated 29 Aug. The console rendered that value as a session id, minted a session-viewer link from it, and printed the same value as a commit three lines below. The link returns 403. The product's own source file states the rule it did not enforce: *"it never invents an id: an absent reference stays absent, because a hold that links to a guessed session is worse than one that links to nothing."

A launch note that told its author to go and spend thirty minutes creating a publication. curl -sI https://morkeeth.substack.com returns 200 and /archive returns 200. A name that really is unclaimed returns 302, which I checked with a control before believing the first result. The publication existed, empty, the whole time. The note was probably right on the day it was written, was never re-run, and was copied into two further documents.

A project board that had been printing no repo at ~/CODE/<name> for a fortnight, for nine repositories that exist, including the one it was itself running out of. git log -S 'async function probeProject' puts the change that broke it at ce63a1c, 16 August 10:18, and the same commit leaves bin/zup.js:307 calling it without await; the fix is ec39f35, 31 August 03:56. Fourteen days and eighteen hours. The standing diagnosis in the repository blamed tilde expansion and named the branch that would fix it. electron/lib/project-state.ts:85 builds its path from os.homedir(), so there is no tilde to expand; the tilde appears only in the error string at shared/timeline.ts:248, which is prose describing where it looked. And the named fix would have been worse than the bug: git diff main..fix/ledger-tilde-expand --stat returns 305 files changed and 30,428 deletions, removing 261 files outright, 48 of which are an entire macOS app. Both halves of a diagnosis nobody had tested, sitting in a file, one merge away from being acted on.

A claim registry whose entire purpose is refusing statements a source does not support. Its shelf header read "120 claims matching" one click from its own counter reading "149 refused." The label filter ran after the SQL limit and then reported the length of the page as the size of the set. The build's regression test now quotes the shipped defect verbatim, which is how I checked it rather than taking the report.

A prompt-grading tool, on the copy published to the package index today. transcripto.py:1447 takes the top five habits; :1448 takes the bottom five, guarded only by if len(rankable) > 5. Worked by hand: at six rankable habits, four rows print under both "SURVIVES MOST, do more of these" and "SURVIVES LEAST, these tend to loop", at the same percentage with the same denominator. At three habits, all three do. At ten, none. A large corpus never sees it. A first-time user is the only person who can.

nobody caught their own

Ten instances. Ten surfaces. Nine repositories. Zero found by the agent that wrote them.

Every one was inside a deliverable whose author had already written a closeout, run its tests, and in most cases rendered the surface and looked at it. These were not rushed builds. The design page shipped with a suite of read-only probes and a browser test that drives the built artifact rather than the source. The competition paper shipped with a verification script a reader could run. The registry shipped with thirteen test functions over the exact page carrying the false header, a number I counted in the file rather than took from a report.

The tests were not weak. They were pointed somewhere else. A test suite grades the code against the code. Not one of these defects is in the code: each is in a sentence about a source, and the source is outside the repository, or in a different file, or on a page on the internet, or in the same file two hundred lines away. There is nothing for a unit test to fail on.

The clearest account of why comes from the person who missed it, in the registry build's own record:

I rendered that page and looked at it and did not see it, because I looked at the default view and not at the filter.

One sentence, and it holds the whole mechanism. Rendering and looking is a real discipline and it caught a great deal else that night, including a stylesheet uppercasing record ids so that a judge copying one would not have matched, which I confirmed against the live payload's lowercase ids. But looking finds what you look at, and you look at what you built. The false sentence is not a rendering fault. It looks correct, because it is a well-formed sentence, written by someone who believed it, sitting next to numbers that are all individually true.

The builder cannot see it for a reason that is structural rather than personal: the builder wrote the sentence and read the source in the same sitting, and after that the sentence is the source. Re-reading your own sentence returns your own memory of the page. It is not a second reading. It is the first one, played back.

what actually worked

Three lanes did catch their own false sentences that night, all three within an hour of fixing someone else's. That looks like the counter-example, and reading what they did makes it the opposite.

The design lane's fix pass installed a check that fails the build if the retracted sentence is ever asserted again, and its first attempt at a second gate was green when it should have been red. The comment its author left in the build script is the clearest sentence anyone wrote that night: it summed "the data the sentence was built from instead of the sentence. Same wrong object as the retracted claim upstairs."

The activity-card lane found that a field on its own brand new code could only ever print zero, because the classifier discarded the records before it asked the question. Right about the wrong object, in the fix for right about the wrong object.

The essay lane wrote "nothing changes places" about its own re-ranking, then re-sorted and found three rows move. I can confirm that one from my own run: on commits only, the highest-surviving habit in the whole table is not the one the essay argues for, it is vague prompts with no named object, at 35.1%, 13 of 37.

None of the three caught themselves by being careful. All three caught themselves while running a procedure that does not care who wrote the thing: re-derive the number from the source, and let a gate fail. The self-catch is not the builder becoming a better reader. It is the builder temporarily stopping being the builder.

Which makes the practical finding small enough to adopt tomorrow. A second pass whose only job is to open every cited source. Not a review, not a critique, not a fresh pair of eyes on the design. One narrow mechanical pass that takes each sentence resting on a source, opens the source, and reads the specific line. It does not require a second person, a second model, or a second machine. It requires that the pass not be allowed to do anything else, because the moment it is also allowed to improve the writing it will start reading the sentence instead of the page.

what I got wrong writing this

The brief I was given said eleven instances in six products. Ten survive the definition, in nine repositories, so the count went down and the spread went up. It also said an inflation of about forty times where I measure about eighty, and thirty percent where the receipt and my re-run both say 29.6%. Those numbers came from a coordinator's summary, which is an authored document, not a measurement. Quoting it back would have been the same defect in a piece about the defect.

Two of my own numbers moved while I was checking them. The record count in the prompt-grading corpus read 412,799 for one lane and 414,213 for me an hour later, because the corpus grows while you read it. My re-derivation of the activity card disagrees with that build's own control by about 10% on the raw total and by eight on the small number. I have printed both rather than pick the one that reads better, because a disagreement between two honest instruments is information and a single confident number is not.

Then I ran the pass on this piece, and it caught four sentences of my own. It said seven of the site's eighteen pieces open with "the" where I had typed eight of twenty, a count I had inherited rather than run. It said the companion file's "four were absence claims" is three, and that I had reached four by listing a universal and a sentence from a product already counted once; the miscount ran in the direction that made the paragraph read better, and they always do. It said a test figure I had reached for was another lane's baseline and not my measurement, so I ran the suite: 562 today over 67 files, of which exactly one test was written the night of the fix.

And it caught one that had nothing to do with numbers. This piece had a section headed "the ask", and the publishing script drops everything from the first heading it recognises onward. "the ask" is on that list. The whole closing section and the sources footnote would have vanished from the published page, silently, with nothing red anywhere, in an essay about surfaces that quietly disagree with their sources. The heading is now "run it on your own logs". I found it by running the script's own regex over the file rather than by reading the file. That is the whole method, in one line.

Then a reviewer who had not written it read the piece, and found three more, all of them in the paragraphs I added after the checking pass had finished. That is the lesson arriving on schedule.

The worst was a clock. Transcript timestamps end in Z, and I read one as Paris wall time, so "three and a half hours before the run started" was really one hour fifty: the quote is 20:04:38Z, which is 22:04 in Paris, against a run that opened at 23:55. It is the same mistake the essay lane made earlier that night with a different quote, in a different file, and I made it after writing that lane's mistake down. I also wrote "he went to bed" straight after a 23:40 quote when my own search output, on the screen at the time, listed four later turns. The real last one is 23:48:41, and one of the night's own builds had independently drawn exactly that moment as the handoff on its card. A corroboration I had read and not used.

The second was the word "shipped". My headline said nine builds each shipped a sentence that night. My own commit dates say four of the sentences were written that night and the other six had been standing for days. The sentence was true of the loudest half and I had generalised it to the whole set: the population error this piece spends a section on, in my own headline.

The third was an absolute in my own frontmatter, "nothing is quoted from another lane's report", contradicted twice in the body by two quotations I put there on purpose. What I meant was about numbers. What I wrote was about everything.

And a limit that eats part of the claim: I checked these ten at their objects, but the map that told me where to look was written by the lanes themselves. A defect nobody reported is not in this count, and there is no reason to think the reported set is complete. Ten is a floor.

and then the headline was wrong

The three above are what the checking pass and one reviewer found. Here is what a second reviewer found, and it is the reason this section is not at the end of the piece by accident.

The count in the headline was wrong. This piece said nine instances. It listed ten. grep -oiE '\bnine\b' | wc -l on that draft returns twenty-four, of which twenty-two name the count and two name something else, so the false number stood in twenty-two places in a draft that lists ten. The section heading over the list read "the other eight" and carried nine bullets. Three different numbers, none of them the list. The reviewer who rejected it counted twenty-one, because a case-sensitive grep does not see the three sentences that begin "Nine"; two instruments, two populations, the same finding. Counting the bullets under that heading takes about ten seconds and I never did it, because I already knew the answer.

The receipt agreed with the list, not with me. receipts/night-as-evidence/14-corrections.txt heads two sections "WHICH OF THE NINE" and then, between them, dates ten: seven authored sentences under one header and three more named under the other. My own evidence file carried the true number and the false headline at the same time, four lines apart, and I wrote both.

The thesis of this piece, arriving inside it. I wrote the count and the list in the same sitting, and after that the count was the list. Re-reading gave me my memory of having counted. It is not a second reading; it is the first one, played back, and that sentence sits three sections above this one, written by me, about someone else.

Going back to the receipts to re-derive the number turned up four more of the same shape, all mine:

And one number in the body was inherited rather than derived: I wrote that the project board had been printing false rows "for four days". Four days is what that repository's own night log says, about a different object, the standing diagnosis. The rows themselves had been printing since 16 August. Fourteen days and eighteen hours. A duration lifted from an authored document and attached to the wrong thing, in the row about a diagnosis lifted from an authored document and attached to the wrong thing.

The honest count is ten: four written during the run, six already standing.

I am not going to end this by saying the fix is to count more carefully. The count was not careless. It was authored, defended, and checked by the person who wrote it, which is exactly the condition the piece says cannot be checked from the inside. It went out wrong and it came back corrected for one reason: it reached someone who had not written the list, and they counted the bullets.

run it on your own logs

If you run agents that write, keep the record they leave, and then read one thing in it: every sentence that names a source. Not the code, not the tests, not the summary. The sentences that make a claim about something outside the repository, and the page each one points at.

The prompt-grading tool used above is one command, offline, on your own logs:

uvx transcripto coach

It will not find this defect for you. Nothing does, yet. But it will show you which of your own instructions left something behind, and the same rule applies to its output as to everything above: when it prints a number, open the thing the number is about before you repeat it.

Numbers: ten instances verified 2026-08-31 between 02:20 and 05:39 CEST, each at its own object, commands in the body. The dates and the four-versus-six split were re-derived from the repositories at 05:39, not taken from this lane's earlier receipt, which mis-pointed four of them. The cited abandonment table was fetched at 02:23 UTC and returned six rows summing to 100%; it is a vendor page whose own methodology says its metrics are modelled, which does not affect this piece, because the claim tested here is only that the row exists. The prompt-grading figures come from re-running the ranking script over 2,917 files and 414,213 records at 04:26 CEST. The two live services were read at 02:25 UTC, with a control in the one case where a missing thing and a present thing return similar-looking results.