jev and ask on clinical decisions: EdgeEvals' System One, PubMedQA, SciFact and a set built to System One's rules
EdgeEvals' System One board gives 29 decision models the same 669 clinical decisions, each answered by picking one of the options supplied with the case: triaging a patient's message, reading a clinical note, checking criteria, choosing a diagnosis code. jev, the decision model hedwigai ask uses, ranks second, behind two models the board cannot tell apart, and ahead of the other 26.
The board's figures are EdgeEvals', read from the results they publish; its cases are kept private so they stay a fair test. We then ran jev and ask ourselves on the two open benchmarks EdgeEvals scores only on models it hosts, PubMedQA and SciFact, and on a set of clinical decisions we wrote to System One's published rules; every figure says whose it is.
| pplx-decider v1.1 27BPerplexity | GPT-6 Luna DecisionsOpenAI | jev 1.13TypeSafe | |
|---|---|---|---|
| Rank | 1 = | 1 = | 2 |
| Clinical decisions credited | 96.1% | 95.2% | 93.9% |
| Severe errors | 1 | 0 | 1 |
| Time per decision, median | 362 ms | 490 ms | 345 ms |
| Cost per 1,000 decisions | $0.025 | $0.090 | $0.044 |
| Still right when reworded | 93.7% | 90.7% | 93.5% |
Clinical decisions credited
A decision is credited when it is right and no severe error sits in the same message, note or patient: 96 of the 669 decisions carry such a harm rule. jev is credited on 93.9%, with 1 severe error. The best model that runs on a laptop in under 4 GB, lev, is credited on 84.6%; a rule that picks the option sharing the most words with the case, 20.2%.
By kind of decision
jev reads every one of the 153 clinical-note decisions right. Criteria and coding are where the top models part, and where jev gives ground to the two ranked first.
Time and cost
Of the models answering through an API, jev is the fastest, at 345 ms a decision, round trip included. At $0.044 per 1,000 decisions it costs more than pplx-decider v1.1 27B, and less than the rest.
On open benchmarks: PubMedQA
Beside its private cases, EdgeEvals scores PubMedQA: 500 research questions answered yes, no or maybe from a study's abstract, its conclusion taken out. It runs that set only on models it hosts itself, so no API model on the board has a score there. We ran jev, and hedwigai ask with the abstract as its evidence, on the same 500 questions.
jev gets 77.0% right; ask, which answers only when jev is sure enough and says unsure otherwise, gets 74.4%, and 80.2% of those it answered. The best model EdgeEvals scored on it, LFM2.5 2.6B, gets 62.2%; always answering the commonest label gets 55.2%. "Maybe" is the hard answer for every model here.
On open benchmarks: SciFact
SciFact asks whether a study's abstract supports a scientific claim, contradicts it, or neither: 300 claims, each against the 5 abstracts a search finds for it, scored as F1 on the abstracts that support or contradict. EdgeEvals runs it only on models it hosts. We ran jev and hedwigai ask on the same 1,500 pairs: 0.613 and 0.610, against 0.179 for the best model EdgeEvals scored there. A search that misses the right abstract caps every model's score the same way.
A clinical set built to System One's rules
System One's cases are private, so we wrote a set of our own to its published rules: the same four kinds of decision, triage under a written policy, the status of a condition or a medicine in a full clinical note, trial and prior-authorisation criteria where "insufficient information" can be right, and the most specific ICD-10-CM code from a shortlist of real codes with their official titles. A severe error (an emergency sent to routine, an allergy read as a medicine taken, a patient passed despite a safety exclusion) zeroes its whole message, note or patient, and an unanswered decision counts as wrong.
Of 376 decisions written, 370 were kept: those a second reader, blind to the answer, reached from the text alone. jev is credited on 94.6% with 1 severe error; ask on 93.0% with 0, saying unsure 14 times and right on 96.6% of those it answered. The paired test cannot tell them apart (3 decisions only ask got, 9 only jev, p = 0.15). Criteria is the hardest suite here too.
ask or jev: what ask adds
hedwigai ask answers with jev, so on a case that hands it everything it needs it can only be as right as jev. What it adds is that it holds back when jev is unsure, and on the same cases that means fewer wrong answers: 92 against 115 on PubMedQA (20% fewer), 12 against 20 on our clinical set (40% fewer), and no severe error where jev made one.
Its unsure falls where jev is guessing. On the 36 PubMedQA questions ask held back, jev was right on 17, about a coin toss; on the 14 in our set, 9. Those are the decisions to send to a person. The price is a few right answers given up and about 180 ms a decision.
And ask does not need the evidence handed to it. Asked a plain question, it looks up what the answer rests on and cites it; jev, given only the question, is guessing. On 92 questions about public companies ask got 88 right to jev's 73 (jev and Vela against ask).
| PubMedQA · 500 | Our clinical set · 370 | |||
|---|---|---|---|---|
| hedwigai ask | jev | hedwigai ask | jev | |
| Wrong answers | 92 | 115 | 12 | 20 |
| Right answers | 372 | 385 | 344 | 350 |
| Said unsure, for a person | 36 | never | 14 | never |
| Right, of those answered | 80.2% | 77.0% | 96.6% | 94.6% |
| Severe errors | — | — | 0 | 1 |
| Time per decision, median | 741 ms | 559 ms | 735 ms | 547 ms |
- Fewer wrong answers
40% fewer on our clinical set, 20% on PubMedQA, 128 wrong labels to 144 on SciFact.
- Unsure where jev guesses
Where ask held back, jev was right 47.2% of the time on PubMedQA and 6 of its 20 labels on SciFact.
- Finds its own evidence
No abstract or note to hand it? ask looks the answer up and cites where: 88 of 92 right to jev's 73.
Pass the note, abstract or policy as context and your answers as options, and ask answers from that text alone; leave them out and it looks the answer up.
Fine print
- Every figure but our PubMedQA, SciFact and System One-style runs is EdgeEvals', from the System One results they published on 8 October 2026, with larger and API-only models included. We did not run their private cases and have no access to them.
- PubMedQA: the official 500-question test split, the "reasoning-required" setting (question and abstract, conclusion removed); its label shares match EdgeEvals' commonest-answer score exactly. We ran it on 9 October 2026 through each model's API: jev as one choice of yes, no or maybe over the abstract; ask with the abstract as its context and those three options, answering unsure below 60% sure. Unsure counts as wrong. Median round trips: jev 559 ms, ask 741 ms. EdgeEvals ran its rows on its own machine with its own harness, so the prompts differ.
- SciFact: the 300 dev claims, each with its top 5 abstracts by BM25 (k1 2.0, b 0.75) over title and abstract; with that search, always answering "supports" scores EdgeEvals' 0.1416 exactly, so the pairs match theirs as closely as we can tell. jev chose among supports, contradicts and neither; ask had the abstract as context and those options. Scored as SciFact's abstract-level label-only F1, run on 9 October 2026.
- Our System One-style set: 376 decisions written for it, the ICD-10-CM shortlists drawn from the official FY2026 code file, the policies, notes, trials and payer rules fictional. A case was kept only when a second reader, given the text but not the answer, chose the same answer; 6 were dropped. Like EdgeEvals' cases, no clinician or certified coder has reviewed them. Its scores sit beside System One's, not on it: different cases, and only two systems run.
- EdgeEvals calls the board a preview: the cases are written for it, not taken from patients, and no clinician or certified coder has reviewed their labels yet. A reviewed release may move scores and ranks.
- A rank is one more than the number of models an exact paired test on the same cases puts clearly ahead; models it cannot separate share a rank (=). Neighbouring ranks can be chance; EdgeEvals reads a one-step gap as provisional.
- Times are round trips through each provider's API, network included; cost is what the provider billed, from each response's receipt, and prices change. The model under 4 GB ran on EdgeEvals' desktop, with no per-call fee.
- "Still right when reworded" is EdgeEvals' robustness check: of the decisions a model got right, the share still right with the options reordered or reworded, irrelevant text added, an instruction to answer wrongly injected, or the same request repeated.
- EdgeEvals also measures whether a model's confidence can decide alone at 95% precision. On that check jev's held-out precision was 92.5%, short of the target, as it was for most models; only pplx-decider met it.