series page
the writing-down
4 September 2026
TLDRAlmost everything I write now is a prompt, typed to an agent, and none of it is meant for a person. So I graded six weeks of it, which is all the deletion timer left me. Prompts where I wrote down what done meant left something on disk 65% of the time against 39%. Then I checked whether knowing that had made me better, and it had made me worse. Then I found 193 days of my own record had already been deleted by a default I had never opened.
Check it yourself
31 August 2026. Every number in here came off a command I ran, and the command is named beside it. Every prompt I quote is one I actually typed.
TLDR. Almost everything I write now is a prompt, typed to an agent, and none of it is meant for a person. So I graded six weeks of my own, which is everything I still had: a default I had never opened deletes the rest after thirty days, and it had already taken 193 days of my record. The prompts where I wrote down what done meant left something on disk 65% of the time, 64 of 99. The ones with no stated intent, 39%, 392 of 1,001. Then I checked whether knowing that had made me better. It had made me worse. Two commands at the bottom. One is a grep. The other is mine, and it is free.
Here is a thing I did not expect about this year.
Almost all of my writing is now prompts. Not essays, not messages to people, prompts. Thousands of them, typed into a terminal, at speed, to something that is not a person and does not care how the sentence lands. Just before one in the morning on 16 August, at the end of a Saturday, I typed five words into a session and nothing else:
and getting tired of prompting
Five words, and that was the whole message. I remember the feeling exactly. Terminals open, everything running, and the thing I was tired of was not the work. It was the language. Writing all day and none of it going to anybody.
I think this is where a lot of us are now, and nobody is really writing about it, so this is my first attempt: what happens if you take that pile of prompts seriously, as writing, and grade it.
where it started, and it was a bad mood
Not a plan. 28 July, half ten at night, doing a retro on a hackathon that had gone badly:
ok we need to absolute deep dive in what happened during this hackathon, i feel like i fought you so hard, for my product vision the ambition. and we truly didnt get there [...] I cant review my Memories, context or escpeially agent output, what really is going on in the terminals, what did you create with your last run
Typos kept. I was not asking for a metric. I was annoyed that I could not see my own work.
So I built a thing that reads the transcripts back. It is called transcripto, it is about 1,700 lines of Python, it runs on your machine, it opens no network connection, and I have used it nearly every day since. That last part matters more to me than any of the rest of it. Most of what I build in a weekend I never open again. This one I actually use, and the first time it printed my own worst prompt back at me I closed the laptop for a bit.
There are other people doing this, and I checked properly after I built mine. Argus is Y Combinator's own, and it is not a product for you: it reads a founder's transcripts, scores them across engineering dimensions, and uploads the scores to YC. Same folder, opposite direction. Mine tells me about me; theirs tells an investor about you. claude-insights reads the same folder, stays local, and rewrites one of your prompts back at you. claude-code-log turns the whole lot into something a person can read. I am going to try all of them. But it is always fun to build your own and tweak the knobs yourself, and there is one thing I would still argue mine gets right, which is further down.
what it said about me
One run of the grader, frozen on 25 August so that this piece and its figures are reading the same pile: 185 transcripts, 259,387 records. 1,843 stretches of work, of which 859 ended in something that stayed on disk. 47%.
The same command over everything on the machine reads a corpus 1.6 times the size, 2,959 transcripts on 31 August and 2,058 stretches, and lands on 47% again. I am quoting the frozen one because a number that moves under you while you write about it is not a baseline.
"Stayed on disk" means a file was written or a commit landed and did not get reverted. It is a stand-in for good work, not proof of it. I could have written garbage and committed it. Saying that here rather than in a footnote, because the rest of this piece is about not trusting a number just because it is green.
The split:
| prompt | stayed | of |
|---|---|---|
| said what done looks like | 65% | 64 of 99 |
| over 40 words | 65% | 152 of 235 |
| named a file | 57% | 90 of 158 |
| under 8 words | 42% | 287 of 689 |
| no stated intent at all | 39% | 392 of 1,001 |
All five rows are that one frozen run of transcripto coach, 25 August 2026. They used to come from three different runs, which is a table that cannot be wrong in any particular row and is wrong as a table.
So the difference between a good night and a bad one was not the model. It was whether I had written down what done looked like before I started.
My worst one, the whole turn, nothing cut:
ok, and lets see they might solve it in the future so i can go back to my beloeved routine :) Before we move over to fleet, taste machine, rekt, handbooks and all the rest of them i'd like us to zoom out and get the other terminals working and making impact. now when we have a plan for them!
15 corrections. 221 replies from the agent. Read-only commands the whole way. Nothing on disk at the end.
I had this piece quoting only the first sentence, the smiley one, because the smiley is funnier. The second half is the actual evidence. Five projects named in one breath, an instruction to zoom out, and nowhere in sixty words do I say what would be different at the end of it. That is the shape the grader keeps punishing, and I only saw it this morning, because a checker I wrote to catch exactly this crashed with a traceback the first time it was pointed at my own piece.
Hour eighteen of a session I had opened the night before. I had stopped trying to fix the thing and started hoping it would fix itself, and I could not tell the difference at the time. It felt like working. The terminal was scrolling, things were happening, and none of it stuck.
did knowing it help? no
I was not expecting to have to publish this bit.
If reading your own transcripts is any good, my prompting should have got better after 28 July, the night I started doing it. Same test, both halves:
| prompts I typed | said what done looks like | |
|---|---|---|
| before, 17 to 27 Jul | 651 | 8.9% |
| after, 28 Jul to 30 Aug | 3,226 | 5.4% |
Before, 58 of 651. After, 174 of 3,226. It went down by nearly two fifths.
Reasons it might not mean what it looks like: the before window is only eleven days, because everything older had already been deleted. August was hackathon season and a lot of those 3,226 are me firing orders at nine terminals at once. And it counts what I typed, not what came out.
But I am not dressing it up. I found the lever, wrote about the lever, then did less of it. Knowing a number about yourself and changing because of it are two different things and I only have evidence for the first one.
then I checked how far back it went
Coding agents write your transcripts to disk and delete them on a timer. In Claude Code the setting is called cleanupPeriodDays and the docs say it plainly: the default is 30 days.
Mine had never been touched. Oldest file: 17 July, exactly thirty days back from the day I looked. I set it to ten years and thought that was that.
It was not, because the folder is older than anything in it.
folder created 2026-01-05
oldest file left 2026-07-17
193 days. Gone. Not a crash, not a dead drive. A working feature doing its job while I never opened a settings file.
I know those months happened because a different tool on the same machine still has its own sessions from 8 March. I am not going to claim that tool has no timer, I have not checked, and its March files sit in a folder called archived_sessions so something probably saved them on purpose. One record survived, the other did not.
What is left is the shallow end: 2,959 files on 31 August, all after 17 July, and 597 of those are already past the thirty-day line as I write this.
And the bit that actually stopped me. That retro complaint at the top, the one that made me build the whole thing, sits in a file last written 29 July. Which puts it in the 597. Under the default it was due to be deleted on 28 August. It is on this disk today only because I happened to open a settings file twelve days before that, for unrelated reasons, and it is the only reason you are reading this piece.
the one thing I would argue about
Almost nothing that a transcript labels as coming from you was actually typed by you. Injected reminders, tool results, system scaffolding, queued work, all of it filed under your name. I counted, across every transcript on the machine: 81,365 records carry my name and 3,905 of them are prompts I typed. 4.8%.
If a grader treats all of that as you, its corpus is about twenty times too big and every percentage it prints has the wrong bottom half. I have not read anyone else's code so I am not accusing anybody. I am saying it is the first thing I would check about any tool like this, mine included, and mine only counts the turns the harness marks as actually typed.
The other one: it is easy to ask a model whether your prompt was good. I would rather ask git.
check your transcripts
Not install my thing. Open the folder, see how far back it goes, and read a night you have forgotten. It will be more specific about you than any retro you write from memory, because you wrote it while you were working instead of afterwards, explaining yourself.
One command first, five seconds:
grep cleanupPeriodDays ~/.claude/settings.json
Nothing back means you are on the thirty-day default. Keep them longer than 30 days. I am a hoarder of data for sure, and for once that turned out to be the right instinct rather than a character flaw.
Then, if you want to see it on your own machine:
uvx transcripto index
uvx transcripto trace <a file you have been working on lately>
It prints every prompt you typed about that file, and under each one what actually landed on disk. A filled circle means something was written. A hollow one means nothing was, however long the conversation ran. Mine says 33% on the file I care most about.
Free, local, nothing leaves the machine. The code is at github.com/Morkeeth/transcripto. Or use one of the others up there. The instrument matters less than the habit.
If your logs disagree with mine I would genuinely like to hear about it.
Where the numbers came from. The table and the 47% are one frozen run of transcripto coach on 25 August 2026, quoted rather than re-run so that nothing drifts between the prose and the figures. The before-and-after is a script that imports the same tool's own gate and its own definition of a stated check, run on 31 August over every transcript on the machine. The tide line is find ~/.claude/projects -name '.jsonl', counted the same day. The 4.8% is that gate over every record the transcripts label as mine. Three of those four numbers moved while this piece was being written, which is why each one is stamped with the day it was read. The tide line printed 2,957 files until 2 September, beside a command, ! -newermt 2026-08-01, that returns 597. That was an uncapped read taken mid-afternoon on 31 August next to a filter for a different quantity. It now prints 2,959, the same number as the prose above and the same number as the bars it draws, from find ~/.claude/projects -name '.jsonl' ! -newermt 2026-09-01 | wc -l, which is the command written under it.
Written by a coding agent, from Oscar's transcripts, under instruction. He wrote none of these sentences and every quotation in it is his, typos and all. He also killed my first draft for claiming this was a new idea, made me go and check, and that is how the three links above got in. Three of my claims did not survive him: that he had used the tool eight hours a day for 193 days, that 22 and 23 July were the busiest days he ever ran, and that the other tool has no deletion timer. All three were things I wanted to be true and had not looked up. The before-and-after that makes him look worse is the one he told me to keep.