HRPulsar
Blog

The question picks the model: Jev and Claude on three of our tasks

· Maxim Berg

Same dog. Different command.

We tested Jev, a new decision model from TypeSafe, against Claude on three tasks from our product. The first is tagging the steps of a business process with the capabilities they need. The second is telling the interviewer from the candidate in a transcript. The third is simple yes/no checks on documents. Claude Opus 5 and Claude Haiku 4.5 also ran on the first task. This is a research note, not a product announcement.

TL;DR

The details

What Jev is

Jev is a model from TypeSafe, released on September 15. It does not write text. You give it a state, which is the context as text or JSON, and a list of typed questions. It returns an answer and a probability for each question. It does not explain its answers. There are three question types (docs). Choice picks one option from a list. Noul gives the probability that a statement is true. Score puts something on an ordered scale. Input costs $0.042 per million tokens, and output is free.

We build HRPulsar, an open-source platform about work and skills. On paper, three of our tasks looked like a good fit for Jev. We ran each task on the same data with Jev and with Claude Sonnet 5, three runs each.

Task 1: which capabilities a process step needs

Our product splits a business process into steps. Then it tags each step with the capabilities it needs, from a list of 17. Examples are checking against a reference, an open-ended conversation with a customer, a judgment about risk, or physical work. One step can need several. The tags decide who can do the step: a person, an AI agent, or nobody yet.

The test set was 10 processes that were not in the prompt examples. That is 104 reference steps, three runs. We fixed the list of steps in advance, so both models tagged the same steps.

Two measures matter here. Precision is the share of right tags among the tags a model set. Recall is the share of needed tags that the model found. F1 combines the two in one number. We needed at least 0.70 for both precision and recall.

For Jev we asked 17 yes/no questions per step, one per capability. The state held the process description and the step. Jev returns a probability for each question, and a tag goes on if the probability is above a threshold. We picked the thresholds on five processes and tested them on the other five. Then we swapped the halves.

Jev found as many right tags as Claude: recall 0.70 against 0.71. But it set about three tags per step. Claude set 1.45, and our reference has 1.62. We count per reference step: the model cut the processes into 113 steps, we matched them to the 104 reference steps, and precision and recall are counted over all tags at once. So Jev's precision was 0.37 against Claude's 0.80, and F1 was 0.48 against 0.75. We also fitted the thresholds on the whole test set, which flatters Jev. F1 still stayed at 0.63–0.64. So the thresholds were not the problem. The probabilities did not separate "needs it" from "does not need it" well enough.

Why did it fail? As far as we can tell, Jev knew about a step roughly what Claude knew. But it could not tell "a tag is missing here" from "this step is complete". We asked it exactly that. Its probability of "nothing is missing" was 0.63 where a tag was missing and 0.62 where nothing was.

Where did the extra tags come from? Most of them were capabilities that the whole process has: a conversation with a customer, bringing people's input together, a judgment about risk. Our guess is that the process description in the state outweighed the step itself. Jev answered "is there a negotiation in this process", not "in this step". We did not test this guess.

One change helped. We replaced the 17 yes/no questions with one Choice over 73 known sets of tags. F1 went up to 0.69. That is still below Claude. It also came from an early run without a pinned model version, so it is a direction, not a result.

Jev was five times faster, 5.6 s per process against 29.8 s. It was about 90 times cheaper, $0.0006 against $0.055. But it stayed below our bar. For a task of this shape it is the wrong tool.

Task 2: who said each line

Some interview transcripts come back as a single voice. Then every line needs a label: interviewer or candidate.

The test set was 31 transcripts with 1,291 lines, three runs. Seven came from our demo. 24 were synthetic, in three languages. They had traps, for example a candidate who asks questions, or short "yeah" lines from both sides.

Jev got one request per transcript. The state held the whole transcript. The questions were one Choice per line: interviewer, candidate or unclear.

Jev got 97.0% of lines right, Sonnet 94.4%. The median time per transcript was 0.63 s against 7–8 s. The cost was $0.0004 against $0.015. That is 12 times faster and 36 times cheaper.

Jev's confidence was useful too. 88% of lines came with confidence 0.95 or higher, and 99.0% of them were right. Below 0.5, two of three were right. You can set a threshold on a number like that.

How we asked mattered more than which model we asked

Before the full run, we tried four ways to ask on one transcript with 44 lines:

Same model, same transcript. Jev could not find a line by its number in a long state. When we quoted the line, it could.

The 1.3 points we owe Claude

In every run, Sonnet failed on one transcript out of 31. The cause was on our side. We limited this call's answer to 4,000 tokens, which is enough for a list of line numbers. But Sonnet 5 thinks before it answers. On a long transcript the thinking used the whole limit, and the call returned nothing. All lines of that transcript then counted as wrong.

Without the failed transcripts, Sonnet was right on 98.4% of lines and Jev on 97.1%. So Claude was 1.3 points ahead. Jev's lead in the headline number came from our settings, not from the model. We went to judge someone else's model and found a bug in our own config. Jev still wins this task on speed and cost. If 1.3 points matter more to you than 12 times the speed at 1/36 of the cost, it does not.

Task 3: three yes/no checks

Before expensive analysis we want three checks. Is this a CV? Is this transcript usable? Is it in the expected language? Each check had 40 cases, half yes and half no, three runs. Jev got one Noul per document. Its thresholds came from the other half of the cases.

Jev let no bad document through and rejected no good one. Sonnet also let nothing bad through. But in every run it rejected the same two good transcripts out of 20. Both were talks about a job offer, and Sonnet decided they were not interviews. Jev took about half a second and $0.00005 per document. Sonnet took 1.3–1.7 seconds and cost about 50 times more.

Other Claude models on task 1

A bigger model did not help on the vague task. Claude Opus 5 set more tags, 1.95 per step. It had the highest recall (0.80) but failed on precision (0.67). We then dropped its tags with its own confidence below 0.6. It tied Sonnet at F1 0.75 but cost 1.8 times more. We picked that cutoff on the same data, so this result flatters Opus. Claude Haiku 4.5 was the cheapest and the fastest Claude model: $0.015 and about 8 s per process. It failed on precision too, at 0.69.

One surprise: Opus was faster than Sonnet here, about 18 s per process against 29 s. It thought in only 14 of 30 calls and wrote about half as many output tokens.

Most of the cost was thinking. By our estimate, thinking was 60–65% of Sonnet's output on this task. With effort: low, the tagging call cost 32% less, $0.0375 instead of $0.055. It also took about 13 s instead of 29. The change in quality stayed inside the noise. A full analysis of one process with Sonnet 5 takes two calls, one to split it into steps and one to tag them, and costs about $0.10.

TaskModelQualityTimeCost
Tag a step with capabilitiesClaude Sonnet 5F1 0.75, precision 0.8029.8 s per process$0.055 per process
Claude Opus 5F1 0.73, precision 0.6717.9 s$0.101
Claude Haiku 4.5F1 0.70, precision 0.698.2 s$0.015
Jev (17 yes/no per step)F1 0.48, precision 0.375.6 s$0.0006
Who said each lineClaude Sonnet 594.4% of lines (98.4% without failed calls)7–8 s per transcript (median)$0.015 per transcript
Jev (one quoted Choice per line)97.0% of lines0.63 s$0.0004
Three yes/no checks, 40 cases eachClaude Sonnet 5nothing bad passed, 2 of 20 good rejected1.3–1.7 s per case$0.0024–0.0032 per case
Jev (one Noul per document)nothing bad passed, nothing good rejectedabout 0.5 s$0.00005

What I take from it

Leaderboards rank models on questions that someone else chose. Our results depended on the shape of our own question. 17 vague yes/no questions about one step, all with the same context, lost. One narrow question per item, with the evidence quoted in the question, won. If a task does not fit "one question, one item, two or three answers", an LLM is still the better tool.

One practical note. Our first run went through Vercel AI Gateway. It did not pin or report the model version, and it rounded probabilities to two decimals. Then we used TypeSafe's own API with jev-1.13.0 pinned. The verdict held, but the thresholds for single capabilities moved by 0.05–0.30. Pin the version, pick thresholds on a held-out half of the data, and pick them again when the version changes.

About the limit on the vague task. Of the extra tags from Opus, about a third broke our rules. A third followed the letter of our rules where our reference had too few tags. The last third fell on rows we still argue about. Our tagging rules lived in five places and contradicted each other. Two taggings of the same kind of process, both made by us, matched on about half the tags. So the limit there was our rules, not the model. That is a separate article.

Caveats

The final Jev numbers come from one pinned version, jev-1.13.0, on one date and on our tasks. Three supporting numbers come from an earlier run through the gateway without a pinned version: the 0.63–0.64 ceiling, the 0.63 against 0.62 for "nothing is missing", and the 0.69 for one Choice over 73 sets. On the pinned version the main result repeated within the run-to-run spread. The no-go on tagging means "wrong tool for this shape of task", not "a bad model". The samples are small: 10 processes, 31 transcripts, 3 × 40 check cases. F1 differences under ±0.05 are noise. 24 of the 31 transcripts are synthetic and were written by Sonnet 5. That is the same model that labelled them, and it may have helped it. No real candidate data went into any model. Everything was anonymized first.

What's the narrowest question in your pipeline that you still pay a reasoning model to answer?

— Maxim Berg, Founder, HRPulsar