hedwigaihedwigaihedwigai
Blog

jev and ask on clinical decisions: EdgeEvals' System One, PubMedQA, SciFact and a set built to System One's rules

EdgeEvals' System One board gives 29 decision models the same 669 clinical decisions, each answered by picking one of the options supplied with the case: triaging a patient's message, reading a clinical note, checking criteria, choosing a diagnosis code. jev, the decision model hedwigai ask uses, ranks second, behind two models the board cannot tell apart, and ahead of the other 26.

The board's figures are EdgeEvals', read from the results they publish; its cases are kept private so they stay a fair test. We then ran jev and ask ourselves on the two open benchmarks EdgeEvals scores only on models it hosts, PubMedQA and SciFact, and on a set of clinical decisions we wrote to System One's published rules; every figure says whose it is.

pplx-decider v1.1 27BPerplexityGPT-6 Luna DecisionsOpenAIjev 1.13TypeSafe
Rank1 =1 =2
Clinical decisions credited96.1%95.2%93.9%
Severe errors101
Time per decision, median362 ms490 ms345 ms
Cost per 1,000 decisions$0.025$0.090$0.044
Still right when reworded93.7%90.7%93.5%

Clinical decisions credited

A decision is credited when it is right and no severe error sits in the same message, note or patient: 96 of the 669 decisions carry such a harm rule. jev is credited on 93.9%, with 1 severe error. The best model that runs on a laptop in under 4 GB, lev, is credited on 84.6%; a rule that picks the option sharing the most words with the case, 20.2%.

1pplx-decider v1.1 27BPerplexity96.1% · 1 severe
2GPT-6 Luna DecisionsOpenAI95.2%
3jev 1.13TypeSafe93.9% · 1 severe
4Clef 27BCloudflare92.8% · 2 severe
5Mistral Large 4Mistral AI92.2%
6d1Liquid AI91.9% · 2 severe
7Solar DecideUpstage90.6%
8Clef-flash 9BCloudflare88.2%
9levInterfaze84.6% · 7 severe

By kind of decision

jev reads every one of the 153 clinical-note decisions right. Criteria and coding are where the top models part, and where jev gives ground to the two ranked first.

Triage
Patient messages: urgency, red flags and where to send them · 240 decisions
1pplx-decider v1.1 27BPerplexity236 of 240
2GPT-6 Luna DecisionsOpenAI234 of 240
3jev 1.13TypeSafe233 of 240
Notes
Clinical notes: whether a condition is present, whether a medicine is taken · 153 decisions
1pplx-decider v1.1 27BPerplexity153 of 153
1jev 1.13TypeSafe153 of 153
3GPT-6 Luna DecisionsOpenAI152 of 153
Criteria
Risk-adjustment support, trial eligibility and prior authorisation · 156 decisions
1GPT-6 Luna DecisionsOpenAI146 of 156
2pplx-decider v1.1 27BPerplexity145 of 156
3jev 1.13TypeSafe140 of 156
Coding
The most specific ICD-10-CM code for a condition · 120 decisions
1pplx-decider v1.1 27BPerplexity109 of 120
2GPT-6 Luna DecisionsOpenAI105 of 120
3jev 1.13TypeSafe102 of 120

Time and cost

Of the models answering through an API, jev is the fastest, at 345 ms a decision, round trip included. At $0.044 per 1,000 decisions it costs more than pplx-decider v1.1 27B, and less than the rest.

Median time per decision, shorter is better
1jev 1.13TypeSafe345 ms
2d1Liquid AI358 ms
3pplx-decider v1.1 27BPerplexity362 ms
4Clef-flash 9BCloudflare443 ms
5GPT-6 Luna DecisionsOpenAI490 ms
6Clef 27BCloudflare630 ms
7Mistral Large 4Mistral AI807 ms
8Solar DecideUpstage1057 ms
Billed per 1,000 decisions, lower is better
1pplx-decider v1.1 27BPerplexity$0.025
2jev 1.13TypeSafe$0.044
3d1Liquid AI$0.046
4Clef-flash 9BCloudflare$0.081
5Solar DecideUpstage$0.087
6GPT-6 Luna DecisionsOpenAI$0.090
7Clef 27BCloudflare$0.217
8Mistral Large 4Mistral AI$0.490

On open benchmarks: PubMedQA

Beside its private cases, EdgeEvals scores PubMedQA: 500 research questions answered yes, no or maybe from a study's abstract, its conclusion taken out. It runs that set only on models it hosts itself, so no API model on the board has a score there. We ran jev, and hedwigai ask with the abstract as its evidence, on the same 500 questions.

jev gets 77.0% right; ask, which answers only when jev is sure enough and says unsure otherwise, gets 74.4%, and 80.2% of those it answered. The best model EdgeEvals scored on it, LFM2.5 2.6B, gets 62.2%; always answering the commonest label gets 55.2%. "Maybe" is the hard answer for every model here.

1jev 1.13TypeSafe77.0% · run by us
2askhedwigai74.4% · 36 unsure · run by us
3LFM2.5 2.6BLiquid AI62.2% · EdgeEvals
4GLiNER2.5-DecideFastino61.0% · answered 463 · EdgeEvals
5Word-overlap rulebaseline56.8% · EdgeEvals
6Most common answerbaseline55.2% · EdgeEvals
7Laya typed-decisionsConvAI Innovations52.2% · EdgeEvals
8GLiClass Instruct EdgeKnowledgator33.2% · EdgeEvals
9DeBERTa v3 Small NLISentence Transformers14.0% · answered 483 · EdgeEvals

On open benchmarks: SciFact

SciFact asks whether a study's abstract supports a scientific claim, contradicts it, or neither: 300 claims, each against the 5 abstracts a search finds for it, scored as F1 on the abstracts that support or contradict. EdgeEvals runs it only on models it hosts. We ran jev and hedwigai ask on the same 1,500 pairs: 0.613 and 0.610, against 0.179 for the best model EdgeEvals scored there. A search that misses the right abstract caps every model's score the same way.

1askhedwigai0.613 · 37 unsure · run by us
2jev 1.13TypeSafe0.610 · run by us
3Laya typed-decisionsConvAI Innovations0.179 · EdgeEvals
4Most common answerbaseline0.142 · EdgeEvals
5GLiNER2.5-DecideFastino0.121 · answered 921 of 1,500 · EdgeEvals

A clinical set built to System One's rules

System One's cases are private, so we wrote a set of our own to its published rules: the same four kinds of decision, triage under a written policy, the status of a condition or a medicine in a full clinical note, trial and prior-authorisation criteria where "insufficient information" can be right, and the most specific ICD-10-CM code from a shortlist of real codes with their official titles. A severe error (an emergency sent to routine, an allergy read as a medicine taken, a patient passed despite a safety exclusion) zeroes its whole message, note or patient, and an unanswered decision counts as wrong.

Of 376 decisions written, 370 were kept: those a second reader, blind to the answer, reached from the text alone. jev is credited on 94.6% with 1 severe error; ask on 93.0% with 0, saying unsure 14 times and right on 96.6% of those it answered. The paired test cannot tell them apart (3 decisions only ask got, 9 only jev, p = 0.15). Criteria is the hardest suite here too.

All 370 decisions
1jev 1.13TypeSafe94.6% · 1 severe
2askhedwigai93.0% · 0 severe · 14 unsure
Triage
118 decisions
1jev 1.13TypeSafe117 of 118
2askhedwigai115 of 118
Notes
96 decisions
1jev 1.13TypeSafe94 of 96
1askhedwigai94 of 96
Criteria
77 decisions
1jev 1.13TypeSafe65 of 77
2askhedwigai63 of 77
Coding
79 decisions
1jev 1.13TypeSafe74 of 79
2askhedwigai72 of 79

ask or jev: what ask adds

hedwigai ask answers with jev, so on a case that hands it everything it needs it can only be as right as jev. What it adds is that it holds back when jev is unsure, and on the same cases that means fewer wrong answers: 92 against 115 on PubMedQA (20% fewer), 12 against 20 on our clinical set (40% fewer), and no severe error where jev made one.

Its unsure falls where jev is guessing. On the 36 PubMedQA questions ask held back, jev was right on 17, about a coin toss; on the 14 in our set, 9. Those are the decisions to send to a person. The price is a few right answers given up and about 180 ms a decision.

And ask does not need the evidence handed to it. Asked a plain question, it looks up what the answer rests on and cites it; jev, given only the question, is guessing. On 92 questions about public companies ask got 88 right to jev's 73 (jev and Vela against ask).

PubMedQA · 500Our clinical set · 370
hedwigai askjevhedwigai askjev
Wrong answers921151220
Right answers372385344350
Said unsure, for a person36never14never
Right, of those answered80.2%77.0%96.6%94.6%
Severe errors——01
Time per decision, median741 ms559 ms735 ms547 ms
  • Fewer wrong answers

    40% fewer on our clinical set, 20% on PubMedQA, 128 wrong labels to 144 on SciFact.

  • Unsure where jev guesses

    Where ask held back, jev was right 47.2% of the time on PubMedQA and 6 of its 20 labels on SciFact.

  • Finds its own evidence

    No abstract or note to hand it? ask looks the answer up and cites where: 88 of 92 right to jev's 73.

Pass the note, abstract or policy as context and your answers as options, and ask answers from that text alone; leave them out and it looks the answer up.

Fine print

  • Every figure but our PubMedQA, SciFact and System One-style runs is EdgeEvals', from the System One results they published on 8 October 2026, with larger and API-only models included. We did not run their private cases and have no access to them.
  • PubMedQA: the official 500-question test split, the "reasoning-required" setting (question and abstract, conclusion removed); its label shares match EdgeEvals' commonest-answer score exactly. We ran it on 9 October 2026 through each model's API: jev as one choice of yes, no or maybe over the abstract; ask with the abstract as its context and those three options, answering unsure below 60% sure. Unsure counts as wrong. Median round trips: jev 559 ms, ask 741 ms. EdgeEvals ran its rows on its own machine with its own harness, so the prompts differ.
  • SciFact: the 300 dev claims, each with its top 5 abstracts by BM25 (k1 2.0, b 0.75) over title and abstract; with that search, always answering "supports" scores EdgeEvals' 0.1416 exactly, so the pairs match theirs as closely as we can tell. jev chose among supports, contradicts and neither; ask had the abstract as context and those options. Scored as SciFact's abstract-level label-only F1, run on 9 October 2026.
  • Our System One-style set: 376 decisions written for it, the ICD-10-CM shortlists drawn from the official FY2026 code file, the policies, notes, trials and payer rules fictional. A case was kept only when a second reader, given the text but not the answer, chose the same answer; 6 were dropped. Like EdgeEvals' cases, no clinician or certified coder has reviewed them. Its scores sit beside System One's, not on it: different cases, and only two systems run.
  • EdgeEvals calls the board a preview: the cases are written for it, not taken from patients, and no clinician or certified coder has reviewed their labels yet. A reviewed release may move scores and ranks.
  • A rank is one more than the number of models an exact paired test on the same cases puts clearly ahead; models it cannot separate share a rank (=). Neighbouring ranks can be chance; EdgeEvals reads a one-step gap as provisional.
  • Times are round trips through each provider's API, network included; cost is what the provider billed, from each response's receipt, and prices change. The model under 4 GB ran on EdgeEvals' desktop, with no per-call fee.
  • "Still right when reworded" is EdgeEvals' robustness check: of the decisions a model got right, the share still right with the options reordered or reworded, irrelevant text added, an instruction to answer wrongly injected, or the same request repeated.
  • EdgeEvals also measures whether a model's confidence can decide alone at 95% precision. On that check jev's held-out precision was 92.5%, short of the target, as it was for most models; only pplx-decider met it.

More head-to-head evals