Oscar Labs

Essays and drafts about building with coding agents. Each one rests on a number out of my own work that you can go and check.

13 pieces. The rough drafts are off the shelf for now.

The finished ones are a series called beloeved: 11 experiments in order, and the next one named with its question.

Reading

Finished enough to argue with. Each one rests on a number out of my own work that you can go and check.

my watch thinks i was resting
I put my Garmin heart rate and sleep over every coding agent session I have on record: 168 million tokens. With agents writing my average heart rate was 64.6, without them 66.2. Sleep after the heaviest days was 11 minutes shorter, which is noise. One person, 81 days, and three numbers I got wrong first.
experiment · 1,397 words · 7 October 2026
a decision model on my own work
Jev said yes to 62 of 63 events. I also tried it on contradictory notes, corrections and my own decisions. Seven graphs, some useful answers, and a much shorter list of jobs I would give it.
experiment · 2,214 words · 8 October 2026
the writing-down
Almost everything I write now is a prompt, typed to an agent, and none of it is meant for a person. So I graded six weeks of it, which is all the deletion timer left me. Prompts where I wrote down what done meant left something on disk 65% of the time against 39%. Then I checked whether knowing that had made me better, and it had made me worse. Then I found 193 days of my own record had already been deleted by a default I had never opened.
series page · 2,355 words · 4 September 2026
row four was always there
One night of coding agents left ten surfaces carrying a sentence their own cited source contradicts, across nine repositories, and not one was caught by the agent that wrote it. In the clearest case a merchant tool announced that no published figure priced an obstacle; the figure was row four of the same table it was already quoting two rows from. Then the piece itself said nine instances over a list of ten, and only a reader who had not written it counted the bullets.
essay · 4,550 words · 31 August 2026
open the page, not your memory of it
The practice without the story: one pass whose only job is to open every cited source, and which is not allowed to do anything else. Four steps, the four shapes that keep recurring, the one-line grep, and the three things that do not work: tests, rendering it and looking, and being careful.
method · 1,550 words · 31 August 2026
nobody edits their prompts
Prompts are the least edited words I produce, and they steer everything. So I graded six weeks of mine. The ones that said what done looks like left something durable behind 65% of the time, 64 of 99. The ones with no stated intent, 39%, 392 of 1,001. On a rerun six days later, counting only real commits instead, it is 30% against 17%. The same command runs on your logs.
essay · 2,579 words · 31 August 2026
my own readme
I gave six agent skills away and spent a week worrying I had sold my edge. Wrong worry. My whole working context lives in markdown I wrote, every model I ask reads those same files, and then they all agree with me.
essay · 1,802 words · 31 August 2026
did i just sell my skills
I put six of my agent skills in a public repo under an MIT licence and handed them to strangers. Then I stopped and asked the obvious thing: what did I think was mine in the first place.
essay · 763 words · 14 August 2026
source of truth for the agent era
Twenty done-claims, each re-derived from the artifact it was a claim about. Eighteen grounded, zero contradicted, two unverifiable. The two it refused to score are the reason the zero means anything.
essay · 1,020 words · 28 August 2026
I gave it away four days before I tried to protect it
Tomorrow I was going to file a provisional patent. This morning I checked where the specification's own reduction to practice lives, and found I had shipped it to a public repo four days earlier. The open-source instinct and the patent instinct were both mine, in the same week, in the same directory, and they never met.
essay · 995 words · 28 August 2026
the wrong object
The guard was right. The number it validated was true, of an object the real system had never seen. A live position ran at 5x cross-margin while the audit log honestly reported leverage 1. Nothing failed, nothing lit red.
essay · 1,264 words · 11 August 2026

Held back

Written, not posted, and the reason is a fact about the world rather than about the writing.

the demo and the number
Two kinds of contest, two ways to lose. A hackathon rewards a demo and a story. A leaderboard rewards a number and does not care who you are. They are not the same game, and the difference is the most useful thing either taught me.
essay · 855 words · 28 August 2026 · not posted
Four of its numbers went stale while it sat. It cites a score ladder topping at 58.05, a leader at 138.25 and a gap of 2.4x. Checked against the live public leaderboard on 28 Aug: the real figures are 68.535, 147.530 and 2.15x. Anyone can falsify it in ten seconds, and the piece argues for not trusting a number that went up. It goes back when it is regenerated from a run rather than hand-edited.
was the competition faulty
I lost a leaderboard competition and went looking for whether the competition itself was broken. It was not, and finding that out was the useful part.
essay · 1,729 words · 28 August 2026 · not posted
Its central number is a leaderboard score. The two submissions this hold originally named have since completed at 7.570 and 61.695, neither near the 68.535 the piece cites, so the reason it was written for expired without anyone noticing. Re-pointed on 28 Aug rather than lifted: three newer submissions are still pending and this hold had never heard of them. A hold whose reason has quietly expired is the one nobody re-reads.