← Oscar Labs

essay

nobody edits their prompts

31 August 2026

TLDRPrompts are the least edited words I produce, and they steer everything. So I graded six weeks of mine. The ones that said what done looks like left something durable behind 65% of the time, 64 of 99. The ones with no stated intent, 39%, 392 of 1,001. On a rerun six days later, counting only real commits instead, it is 30% against 17%. The same command runs on your logs.

Check it yourself

I ran coding agents for months before it occurred to me that the thing I type into them is writing, and that all of it was sitting on my disk in a folder I had never opened.

Then I opened it. 2,957 transcript files, and the folder is 193 days older than anything in it. It was created on 5 January; the oldest thing left is 17 July, which was exactly thirty days back from the day I looked. A default I had never opened had been sweeping the rest. So six weeks is not a choice, it is everything the timer left me.

So I graded it. Prompts that stated what done looks like left something durable behind 65% of the time, 64 of 99. Prompts with no stated intent, 39%, 392 of 1,001. The difference between my good nights and my bad ones was not the model, it was whether I wrote the check.

Below is the grading, the caveat that eats most of it, and the one command that runs it on your logs instead of mine.

A coding agent is a tool you type an instruction into and then walk away from while it edits your files, runs commands and writes code for the next forty minutes on its own. The instruction is the whole steering wheel. And we edit our essays. We edit our launch posts and our commit messages and our cover letters. Then we open a terminal and type "fix the thing" to one of these systems and we do not read that sentence back once.

So I graded them.

I ran a grader over 185 of my own coding-agent transcripts, as they stood on 25 August 2026. 259,387 records, most of it machine noise. What survived a filter for genuine human turns was 3,438 prompts, ranked into 1,843 episodes. 859 of those episodes survived.

Survived is a proxy, and a generous one. It means the episode reached a durable write or an un-reverted commit, so one saved file counts the same as a commit that landed and stayed. It does not mean the work was correct, or shipped, or worth keeping. That caveat travels with every number here, because the piece is partly about refusing to trust a green number, and a piece like that does not get to launder its own.

So here is the strict cut, before the table rather than after it. Throw the saved files away and count only un-reverted commits. Prompts that state a check land one 30% of the time, 32 of 108. Prompts with no stated intent, 17%, 186 of 1,101. Every bucket over a hundred episodes lands between 42 and 48% of its own survival rate, so the whole table roughly halves. The ordering this piece rests on holds under both definitions: states-a-check, then terse, then no stated intent, in that order either way. The ratio between the top and the bottom of those three widens, 1.6x to 1.8x, while the distance in points halves along with everything else. The rows do shuffle: two pairs that were already within a point and a half of each other swap, and one row moves a long way. That one is the row that argues with me, so it gets its own paragraph further down rather than a footnote here.

That cut is on tonight's corpus, not on the frozen August one, and the reason is worth saying out loud: the August report survives only as a text file, so it can give me per-bucket survival and not per-bucket commits. Corpus-wide it can do both, and it agrees. 342 of its 1,843 episodes reached a commit, 18.6%, against 47% that merely stayed on disk.

Which number is the real one depends on what you think a proxy is for, and I am not going to pretend that is settled. 65% is what the tool prints. 30% is what the word commit means. I have quoted the first one on five surfaces for six days without noticing they were different, in a piece whose entire argument is that a confident sentence should be opened at its source.

Here is the top of my corpus.

BELOEVED  /  01Which of my prompts survivedSolid survived. Hollow did not. One bar is blue, and it is the loss.states a check or done-condition65%64 / 99terse, under eight words42%287 / 689no stated intent at all609 left nothing39%392 / 1,001Survival = a durable write or an un-reverted commit. A proxy, not proof the work was correct.TRANSCRIPTO COACH  /  185 TRANSCRIPTS  /  FROZEN 2026-08-25
01-survival SOURCE · regenerated by build-figures.py · 2026-08-31

Prompts that state a check or a done-condition survive 65% of the time. 64 of 99. Terse prompts under eight words survive 42%. 287 of 689. Prompts with no stated intent at all survive 39%. 392 of 1,001.

Nothing in that table is about the model.

The discourse about coding agents points the other way, almost always. The model failed me. The agent looped. The context window ate my plan. I have said all three this month. Then I pointed the grader at myself.

My best landed prompt in the sample is not clever. It names a terminal, a repo path, and a job. Zero corrections, commit witnessed.

Here is a check sentence out of another one, typed on 11 August:

Every fix needs a test that is red today and green after.

Twelve words. Nothing about the model, nothing clever, no technique. It just makes it impossible for either of us to call the thing done while it is not.

My worst reads like a shrug, and for four drafts I quoted only the shrug:

ok, and lets see they might solve it in the future so i can go back to my beloeved routine :)

That is the first sentence of it. Here is the whole typed turn:

ok, and lets see they might solve it in the future so i can go back to my beloeved routine :) Before we move over to fleet, taste machine, rekt, handbooks and all the rest of them i'd like us to zoom out and get the other terminals working and making impact. now when we have a plan for them!

Fifteen corrections. Two hundred and twenty-one assistant turns. Read-only bash. Nothing written down that stayed.

I cut it at the smiley because the smiley made the point I wanted, and the second half is the point. It names five projects. It asks for a specific move. That is an intent, and I had filed it under prompts with none. What it never does is say what done would look like, because "making impact" cannot be graded, and neither of us could have told you at the end of it whether we got there. A goal is not a check. This is the clearest example of the difference in the whole corpus and I had it quoted in half.

I am not proud of the shrug. I am putting it here because the data is not flattering, and flattering data would not be useful. The industry story is that the agent failed the human. On my machine, whether I wrote the check was as strong a survival signal as anything else in the corpus, and it is the only one I control by typing.

I know exactly what I was doing when I typed it. The session had been open for eighteen hours and sixteen minutes. That sentence was the hundred and fifth thing I typed into it. I had stopped trying to fix the thing and started hoping it would fix itself, and the smiley on the end is where the hoping shows. Two hundred and twenty-one assistant turns later there was nothing on disk. Not a bad commit. Nothing.

The uncomfortable part is not that the prompt was bad. It is that I could not have told you at the time. It felt like working. I was typing, the agent was replying, the terminal was scrolling, and it looked exactly like the nights that produced something. The only way I found out was by grading six weeks of them at once, and the difference between the good nights and that one was not effort, or the model, or the hour. It was whether I had written down what done meant before I started.

People who write about writing tend to give up at the same place. On finding the right words, the honest answer is that taste does not reduce to rules, and everyone says so eventually.

Prompts are a different genre with a much lower bar, and that is exactly why the bar is worth measuring. An essay has to be good. A prompt only has to be unambiguous about what done means. That is a smaller target, and unlike taste, you can count whether you hit it.

So this is not "write better prompts", which is already a genre and most of it is costume. It is narrower and meaner. On this operator, on this harness, over this window, the quality of the writing in the instruction predicted whether anything durable happened next. Detailed prompts over forty words scored level with prompts that carried a check. Citing a file path helped, by less. Having no intent at all hurt.

The honest limits, in one place. One operator, one seat. Survival is a proxy for durability, not for truth. I have not run a controlled experiment isolating "stated a done-condition" from "was long" or from "I already knew what I wanted". The report that produced these numbers says so on line one of its method. I am not laundering a personal diary into a law of nature.

What I will claim is smaller. When I skip the writing-down, my own agents waste my own time at a rate I can now measure. When I put the check in the prompt, the episode is more likely to leave a commit. That is enough to change how I work tomorrow morning.

Here is the part that matters more than my number.

The grader is yours. One command, it runs on your own logs, it never opens a socket, and it prints your table instead of mine.

uvx transcripto coach

The table above is frozen at 25 August. I ran the same command again tonight, against everything on this disk, to see whether the shape held.

YOUR PROMPT HABITS, GRADED   (offline, your machine only)

SURVIVAL IS A PROXY: survival = a durable Write/Edit or an un-reverted git commit in-episode. A PROXY, not proof the work was correct or shipped.

SURVIVES MOST do more of these: 65% (71/110) states-a-check-or-done-condition 64% (178/276) detailed (>40 words) 59% (22/37) no-object (pronoun/vague) 59% (451/767) intent:CHANGE 57% (106/187) cites-a-file-or-path

SURVIVES LEAST these tend to loop: 31% (4/13) intent:REVERT 40% (451/1135) intent:none 42% (320/771) terse (<8 words) 42% (81/193) intent:DESCRIBE 45% (5/11) intent:TEST

harness claude · 2957 transcript(s), 431,361 records · 3981 typed by you (0.92%) episodes: 2119 ranked, 992 survived (47%) commit 454 · write/edit 538 · reverted 2 · nothing durable 1125

───────────────────────────────────────────────────────── compare yours. numbers only, nothing from your prompts:

transcripto coach · claude · 2119 episodes · 47% survived states-a-check-or-done-condition 65% (71/110) intent:none 40% (451/1135) gap 1.6x ───────────────────────────────────────────────────────── ```

SOURCE: uvx transcripto coach on PyPI 0.1.3, run 2026-08-31, captured verbatim to receipts/coach-uvx-0.1.3-2026-08-31.txt. 0.1.3 shipped the same day and changed the layout: the worst-prompt reveal moved to the top and a paste-able summary was added at the bottom, so the 0.1.1 block this piece carried until tonight was no longer what the published command prints. Every ranked row the tool printed is here. Nothing is trimmed. An earlier draft cut three rows "for width", and one of the three was the third-highest row in the table. The best-prompt line is removed because it names a private repo path. Nothing else changed.

64%, 69 of 108, against 65%, 64 of 99 six days earlier. Same worst prompt, still at the bottom.

Now the row that argues with me. Prompts with no concrete object in them, the pronouns, the "fix it", come third at 59%, above citing a file path. They were third in August too, at 60%. Under the strict cut they are not third, they are first, at 35%, above everything. It is the only row that moves a long way when you change the definition, and it moves the wrong way for me. It is also one of the smallest buckets in the table, 37 episodes, and I am not building a story on 37 either way. I am leaving it printed because it was in my output, and cutting it was the tidier lie.

Two caveats on that header, and they are the exact kind this piece is about. The first: the tool's own transcript counter reads 185 in the August report and 2,910 tonight, while records moved only from 259,387 to 412,828. Those two counts are not the same unit, so I am not going to tell you the corpus grew fifteenfold. The second: I ran the grader five times across about an hour while editing this paragraph, on a machine I was not otherwise using, and the record count came back 412,645, then 412,772, then 412,799, then 412,828, then 413,089. Every ranked row was identical every time, and so was the episode count, 2,058. The last of those five is two different entry points, the tool and the script I use to re-cut it, run back to back: they agreed with each other exactly. So the counter really is moving under me rather than two instruments disagreeing, and the table is stable while the population line is not. The population line is a number that is correct about the wrong object, sitting in my own instrument, in the piece where I tell you not to trust one.

Run it and find yours. Same proxy, same caveat, your corpus. If it ever prints something that flatters you, distrust it, because a generous proxy is a broken one.

I would rather hand over the instrument than the finding. A finding you can only take my word for is an anecdote with a percentage sign on it.

For an essay, whether the writing is doing its job needs an editor and a lot of taste. For a prompt, the answer is sitting in a log file on your own disk, and you have never opened it.

I used to think I was debugging the agent. Half the time I was debugging the blank where the done-condition should have been.

Numbers: my own coach report, 185 transcripts, 2026-08-25, frozen, kept at receipts/coach-frozen-2026-08-25.txt. The rerun is the same command on 2026-08-31, kept at receipts/coach-uvx-0.1.3-2026-08-31.txt. Survival means a durable write or an un-reverted commit, which is a proxy and a generous one. The strict cut, commits only, is 30% (32 of 108) against 17% (186 of 1,101) on the 31 August corpus; the shipped tool does not print it, so the script that derives it is at receipts/strict-cut.py and its output at receipts/strict-cut-2026-08-31.txt. The shrug is timestamped 2026-08-14T16:51:36Z, 18 hours 16 minutes after that session's first record and the 105th of 152 prompts I typed into it, all four figures read off the transcript rather than off my memory of the night. Run your own: uvx transcripto coach. Source: github.com/Morkeeth/transcripto.