← Oscar Labs

beloeved

Experiments on my own work

One number, one command you can run, and the half that makes me look worse. Every piece takes a confident result about my own work, goes back to the object the claim was about, and publishes the gap. Every number carries its denominator. Every piece ends in something you can paste, so you get your number instead of taking mine.

The experiments, in order

11 published. Each one rests on a number out of my own work that you can go and check. The order is the reading order, not a chronology.

01
the writing-down
I graded six weeks of my own prompts, which is all a deletion default I had never opened had left me. The ones that wrote down what done meant survived 65% of the time, 64 of 99. The ones that said nothing, 39%, 392 of 1,001.
02
row four was always there
One night of coding agents left ten surfaces asserting a sentence their own cited source contradicts, across nine repositories. Not one was caught by the agent that wrote it.
03
open the page, not your memory of it
The practice with the story taken out: one pass whose only job is to open every cited source, and which is not allowed to do anything else.
04
nobody edits their prompts
The same grading at full length, then rerun six days later counting only real commits instead of any durable write: 30%, 32 of 108, against 17%, 186 of 1,101. The ordering holds under both definitions.
05
my own readme
I gave six agent skills away and spent a week thinking I had sold my edge. Every model I ask reads the same markdown I wrote, so they all agree with me.
06
did i just sell my skills
The short version of that week, and the question underneath it: what did I think was mine in the first place.
07
source of truth for the agent era
Twenty done-claims, each re-derived from the artifact it was a claim about. Eighteen grounded, zero contradicted, two refused, and the two refusals are why the zero counts.
08
I gave it away four days before I tried to protect it
The day before filing a provisional patent I checked where its own reduction to practice lived, and found I had shipped it to a public repo four days earlier.
09
the wrong object
The guard was right. A live position ran at 5x cross-margin while the audit log honestly reported leverage 1. Nothing failed and nothing lit red.
10
a decision model on my own work
Jev included 62 of 63 events and matched 17 earlier brief choices. Seven graphs show what happened across notes, corrections and my own decisions.
11
my watch thinks i was resting
Garmin heart rate over 168 million agent tokens. 64.6 beats per minute with agents writing (639 hours), 66.2 without (279 hours). Sleep moved 11 minutes over 68 nights.

Next

Its number is hollow because the experiment is not on the record yet. That is the same rule every figure in the series uses: solid survived, hollow did not.

Get the next experiment by email

One piece at a time, no cadence I am going to pretend to keep. If your own logs disagree with mine I would genuinely like to hear about it.