experiment
a decision model on my own work
8 October 2026
My daily brief is supposed to help me find things worth leaving the house for. On 2 October, I tried asking Jev to make one of its decisions: should this event or hackathon go in on my list?
The candidates included rock-paper-scissors by the Eiffel Tower, AI Builders and Databases Paris, and Dog fest. Different dates, all real listings from the history of my brief.
There were 63 candidates in the test set. The version of Jev selected in an earlier round said yes to 62 of them.
Helpful If I would have a longer week.

It agreed with the earlier brief on 17 of the 63. A one-line rule, include it if it is a Luma event, agreed on 53.
That was not what I had seen in the memory test. There, Jev had matched the model I already used. Same model, another job, a very different result.
So this became a bigger experiment. Could it check my notes? Tell whether a source supported a claim? Notice when I was correcting something? We had also tried asking it to predict decisions I had already made.
There are plenty of small choices I would like help with. Apparently we needed to be more specific about which ones.
first, what is Jev?
Jev is TypeSafe’s model for decisions with a fixed set of possible answers. You give it the information, explain what you are deciding, and ask for a yes/no probability, a choice, or a score.
I tested Jev 1.13. The attraction is easy to understand. If I need to know whether two notes disagree, I do not need a paragraph every time. I need an answer I can use.
The memory test had 26 pairs of notes: 14 based on facts I had previously settled, and 12 harder comparisons. Jev got 25 right. So did Qwen 3.6 Plus. That result repeated across three runs.
Both missed the same pair, which compared a number of devices with a number of users. I am not convinced the test’s answer was right. Those can be different populations. Even the neat tie comes with a question about the test.
Still, Jev’s median response time was about 0.3 seconds, against Qwen’s 14.2. The recorded cost per answer was about 105 times lower. That is for these particular questions and settings, but it is a good reason to look further.
choosing is harder than saying yes
The brief has limited space. Choosing one thing depends on what else is available and what I tend to find interesting.
Our setup asked about each candidate separately. We chose how to ask on 75 candidates, then tested it on the 63 held back. That is where it agreed with the earlier brief on 17. Re-running the earlier brief’s prompt agreed on 45. The simple Luma rule agreed on 53.
Rock-paper-scissors shows how much the question mattered. Asked to choose “include” or “drop”, Jev said drop. Asked for a yes/no probability, it returned 0.14. The cutoff selected in the earlier round was 0.10, so that version included it.
Same event, different ways of asking, different answers to the thing I needed to decide.
These scores do not measure good taste. The answers we compared against came from what another model had included in earlier briefs. I had not personally rated all 63 candidates. The brief’s prompt had changed during that period too.
I did not prove that Jev cannot select events. This version failed to reproduce the selections my brief had made. It also left the interesting question unanswered: which few things would I actually want to do?
a very confident wrong answer
The source-checking experiment looked good at first. Given a claim and text from a document, does the text support it?
Jev matched 28 of 29 answers in the test. Gemini matched 27, and a string-matching check matched 22. The answers were written by other models, with 23 of the 29 examples coming from one regulation. A one-point lead is a small thing here.
The Jev mistake is more interesting. It accepted evidence from the wrong annex. The probability it gave to “supports” was 0.99.
A very confident answer from the wrong part of the document.
That changes how I would use it. I would want to check the exact cited section as well as whether the text sounds supportive. Letting the confident answers through would have let this one through.
can it tell when I am saying no?
Another experiment looked for corrections in my conversations with AI. If I have already said something is wrong, it would be useful to notice that and keep the correction.
There were 33 messages labelled as corrections among the 92 held back for this test. Jev found 25 of them and incorrectly flagged seven other messages. Haiku found 26, with six false alarms. The keyword rule found 21, with eight false alarms.

Jev’s recorded answers cost roughly 16 times less than Haiku’s. One fewer correction found, one more false alarm, much cheaper to repeat. That is a trade worth looking at.
The labels came from one model rater, and the keyword rule had already been tuned on these messages, which favours it in this comparison. Before I trust something to remember what I meant, I would like to check those answers myself.
We also tried checking claims that a piece of work was finished. Did the attached evidence actually support that claim? In each run of 28 written test cases, the system made 21 correct definite calls, left five for a person and withheld two for privacy.
No wrong definite calls in those runs. But all five cases left for a person were actually labelled supported. Being cautious meant more checking for me. These were constructed examples, not a count of finished work from my day.
then we asked what I would do
The September experiments were more personal. We took decisions I had made about my projects and asked Jev to predict them.
One test had 22 decisions. Seventeen came with a recommendation from the AI. We asked the same questions with that recommendation hidden, then visible.
With it hidden, Jev picked the recommended option on seven of those 17 cases. With it visible, that became 15 of 17.
I had rejected the recommendation in 12 cases. Jev predicted eight of those 12 when it could not see the recommendation. Once it could see it, that fell to two.

So, more agreement with the AI. Less agreement with me.
I would want a second opinion to have a chance of disagreeing. This is a reason to ask it before showing it the first answer.
We also gave it a written profile of how I make decisions. On 16 later decisions, that version matched 13, compared with five for the plain version.
Then we tried it on the earlier 22. The profile version matched ten. The plain version had matched twelve.

The examples were prepared with the outcomes already known, and the profile overlapped with some of those decisions. These are small retrospective tests. They cannot establish that the model learned my taste.
Even with that helpful setup, the improvement did not carry across the two sets. A description of me was useful on one set and less useful on the other. I would not use that to decide when it was safe to stop asking me.
the small print became the next test
TypeSafe documents weaknesses with arithmetic and dates. We tested those directly, using written examples with answers calculated in code.
There were 248 pairs, each tested twice. Jev got 459 of the 496 judgments right. All 37 errors were missed contradictions. It raised no false alarms.
Then you look inside that number.
It caught none of the ten wrong-sum judgments. It also missed all twelve cases where a written weekday disagreed with a date.

The wrong sums did not make it noticeably unsure. It gave them almost the same probabilities as the correct sums.
So we tried doing the arithmetic and date checks in code, then using Jev for the remaining judgment and Qwen for uncertain cases. That combination reached 494 of 496.

The checks were developed against those examples. A differently worded set also helped us refine them, so neither is a blind test of the fix. On that second set, the combination scored 84 of 92, against Jev’s 74 and Qwen’s 90. One weekday label is disputed too.
There was another less flattering result. The saved scan of my real notes reported that the arithmetic checks settled none of 67,747 candidate pairs. We had made the written tests better. Those particular checks had not yet helped with the notes I actually had.
and if I ask again?
We repeated some identical questions. One borderline answer moved between 0.48 and 0.61 across five calls. If I make a decision at 0.50, that gives me different decisions from the same information.

A clearer example stayed between 0.89 and 0.91. The problem shows up around the point where I turn a number into a yes or a no.
We also asked it to choose among 41 documentation pages. In each of ten runs, 36 or 37 pages received exactly zero. The top choice stayed the same, but a fifth place on the shortlist could come down to a tie between zeros.
Asking a separate yes/no question about each page produced more distinct scores and the same top-five set across three runs. That was one query. It tells me something about how to ask for a shortlist, rather than proving that the shortlist was good.
Even asking several questions together changed answers. In a small correction test, one message went from 0.88 on its own to 0.32 in a group of ten. The saving was only about 1.6 times on that run. I would rather keep the questions separate until I understand that better.
next: one evening, three options
For the next test, I want this to end with me actually going somewhere.
Three events I could realistically attend on the same evening. I pick one and write down why before seeing Jev’s answer. Jev gets the same options, the date, the travel involved, what it costs and a short description of what I tend to enjoy. It has to choose one, with staying home allowed too.
Then we compare. Did it pick what I picked? Did reading its choice change my mind? Where did I actually go, and was I glad I went?
Those are different questions. A model could disagree with my first choice and suggest something I end up enjoying. Or it could predict my choice perfectly and we both pick a disappointing evening. If I never go because it rains, that belongs in the result too.
I would also like to repeat the personal-decision test on choices I have not made yet. Keep the profile fixed, record the prediction, then make the decision. No writing the test while already knowing my answer.
That gives the next part something this one does not have: fresh choices, my own ratings, and a result outside the screen.
Jev is cool, because there are plenty of small decisions in my tools. These runs gave me a more specific place to try it. They also gave me a shortlist of jobs I would keep out of its way.
Now I want to see whether it can make a shorter shortlist for me.
Methods: Jev 1.13, hosted version typesafe/jev-1.13-20260917. Main experiments: 2 October 2026. Personal-decision tests: 23 to 24 September. Memory: 26 pairs, three runs. Source support: 29 machine-labelled items. Corrections: 92 held-out messages with machine-made labels. Brief: 75 development and 63 held-out candidates, compared with earlier model choices. Finished-work checks: 28 constructed cases, three runs. Stress: 248 pairs, two runs; both follow-up sets informed development. Personal choices: 22 and 16 retrospective decisions, with author-knowledge and profile-overlap limits. Ranking: one query, 41 pages. Repeat and grouped-question tests were small diagnostics. Costs and times describe the recorded runs. The proposed evening test has not happened.
How this was made: AI helped run the experiments and prepare this draft. Numbers were checked against the saved results. The datasets and full run records are not published with it. This is a personal investigation, not an independent benchmark.