← Oscar Labs

essay

my own readme

31 August 2026

TLDRI gave six agent skills away and spent a week worrying I had sold my edge. Wrong worry. My whole working context lives in markdown I wrote, every model I ask reads those same files, and then they all agree with me.

I gave six of my agent skills away as markdown files and spent a week worrying I had sold my edge. The worry was the wrong one. The real problem is that my whole working context lives in markdown files I wrote, every model I ask reads those same files, and then they all agree with me. That agreement is not evidence. It is my own notes read back to me in a different voice.

If that setup is foreign, here is the shape of it. I do not keep my projects' state in a database or a project-management tool. I keep it in plain text. One file describing how each repo works, one holding each plan, one note per idea, one file per procedure. An agent opens a session, reads that pile, and everything after that is downstream of it.

I have a name for the failure and I want to hand the name over, because the fix is smaller than the problem: state your core problem in one sentence with zero documents attached, and get one answer from that sentence alone, before anybody gets to read your docs.

Here is the one number, early, with what it is out of.

Same question, two amounts of my notes attached nothing but the problem sentence 0% housekeeping, not new ideas my whole repo attached 64% housekeeping, not new ideas Pilot, n=2 problems, one model. Not a sample. The direction, not the cause.
the context pilot, n=2 SOURCE · context pilot, 2 problems, briefs/2026-08-19

I put the same design question to a model twice. Asked cold, with nothing but the problem sentence, none of what came back was housekeeping and it proposed a different product. Asked with my whole repo attached, 64% of what came back was housekeeping. Two problems, so that is a pilot and not a sample, and a good half of this piece is about why I cannot make it more than that.

First, what it looks like when the pile turns on you.

The loud version is a session dropping its context mid-work. This is me, mid August, typos kept:

Wrong names of the terminals, big issue and a lot of confusion Where did the handbook lands Taste machine, where are the tweets? Houve missed so many terminals, we loose it honestly This life is so weird, im so out of it 1.18 and my terminal lost all context

Everybody complains about that one. It is the failure at its loudest. The quiet version is worse, because the pile does not vanish. It agrees with you.

Here is the quiet one, with a timestamp, because it is the moment I found the name.

On 10 August an agent had spent a night helping me think about how people actually get hired. It kept citing one figure at me, roughly 0.4% one way against roughly 40% the other, a hundredfold gap, and it built most of an argument on top of it. Then it stopped and wrote this:

It came from your own README and I never checked it.

And then:

the number in your README is inflated by about an order of magnitude, and I repeated it as fact.

A minute later I typed:

hahah my own readme, this is the ENSEMBLE theater once again, reading my broken context and findings like truths.

Ensemble theater is what happens when everyone in the room has read the same documents, the room converges, and the convergence feels like a finding. Nobody is out of tune. They are all playing off one score, I wrote the score, and parts of it are wrong.

Note what did not happen. Nothing failed. No error, no crash, no lost context. A wrong number I had written down myself was read back to me all night in a confident voice by the thing I had asked to check things, and the only reason it ever came out is that the agent went and looked at the primary source on its own. If it had not, I would still be quoting a hundredfold.

One night holds the whole problem. When two files disagree, or one is nine days stale, the agent cannot tell. It reads a stale sentence with exactly the confidence it reads a true one. Then I ask a second model, and the second model reads the same pile. Two opinions, one source.

I said this out loud on 18 August without quite meaning to:

Also i think one misstake might be feeding TOO much context thats how we get the ensemble theater ofc all the MD files tells the project makes sense, I might just need a fresh pair of eyes of the problems im trying to solve, do that too please.

So we tried to measure it. Same design question at three levels of context: the raw problem sentence alone, the problem plus one paragraph of what I wanted, and the problem plus the entire repo, docs and code. Then every recommendation went into one of two piles. Changes to what the product actually is, versus code-level housekeeping. Wire this, delete that dead file, rename this field.

The cold runs, given nothing but the problem sentence, came back 0% housekeeping on 2 of 2 problems, zero housekeeping moves out of six recommendations each time, and both proposed a different central idea for the product. The full-context runs from one model came back 7 of 9 housekeeping on the first problem and 4 of 8 on the second, 64% on average, and proposed no new central idea on either.

Which is the shape I expected, and it is where I have to stop, because the run is not clean. Two problems is a pilot, not a sample. The context level was perfectly tangled up with which model ran it. And a second model at full context came back 14% housekeeping and stayed as radical as the cold pass. The direction showed up. The cause did not. It is registered as unmeasurable until someone runs the crossed version, and I would rather say that than sell you a law.

What survives the caveats is a symptom I can spot in ten seconds, and I wrote that down too, the same day:

symptomatic for ensemble therater is plumbing and incremental improvement.

When a review comes back full of small correct fixes and nothing about what the thing is for, that is not a healthy codebase. That is a room that read my notes.

Now the skills part, because it is the same object.

A skill is a markdown file. That is the whole technology, and I liked that fact enough to name my competition entry out of it, on 12 August, in a list of candidates I was typing at an agent:

A skill issue, let the best MD file win ( THIS ONE IS GOOD NICE)

Six of mine are MIT-licensed markdown, 30,846 bytes of text, and a stranger can clone them and run them. The competition was about whether skills like that actually help. My entry could not show that they help. The benchmark mounted my six skills on 8 tasks and a skill body actually opened on 3 of those 8, so any lift number was an average over runs where the skill was never in the room. I reported the null.

Which sounds like a loss and is the reason I stopped worrying. If I cannot demonstrate the six files lift anything, I did not sell my edge. I sold some formatting.

Take the other branch seriously though, because the flip answer is too easy. Suppose the skills did work. Suppose each of the six carried real lift. Then a skill file is a commodity by design. It is meant to be opened by a stranger and do the same thing for them that it does for me. The entire point of writing a procedure down is that the writing down travels. If the value were in the files, giving them away would be exactly the mistake the title is afraid of.

But that is not where the value turned out to sit. The thing I would defend is not in any of the six files. It is one sentence written under my name in a contest where every incentive points at claiming a number: we cannot show these six skills lift an agent.

None of that is in the six files. It is in the refusal to trust the green number. That refusal is the thing that did not package.

Which would be a clean ending, except it is not, and I know it is not.

The refusal is also a procedure, and I have already written it down as a file. I keep one that asks what would prove this is right before I build it, and another that asks whether a stranger could run it at all. The part I called mine is the part I am busiest turning into markdown a stranger can run. So the recursion closes on me. If I can write down how I distrust a number, that becomes a file too, and then I have sold that as well.

There is no command at the end of this one. Everything else I have written this month hands you a script and tells you to distrust your own number. This one hands you a sentence and a minute, because a gate against agreeing with yourself is not a thing I can package, and bolting a tool onto it would be the same mistake in a new costume.

So here is the minute. Before your next build or review, write your core problem in one sentence, with zero documents attached, and get one answer from that sentence alone. If a model with none of your context builds the right thing from that line, that line is your product's reason to exist. If it builds something different from what your full-context room agreed on, your room was reading your notes.

I sold the skills that were always going to be sellable, and I am now busy packaging the one I told myself was mine. The uncomfortable part is not that I gave six files away. It is that I can no longer point at the thing I would not.

Numbers: 6 skills, 30,846 bytes of markdown measured 2026-08-31, MIT. The benchmark mounted them on 8 tasks and a skill body opened on 3. Context pilot: 2 problems, 3 context levels, scored un-blind by its author, context confounded with model, verdict unmeasurable until someone runs the crossed version. Every quoted line is my own typed turn, checked against the record rather than against my memory of it.