Self-hosted decisions for ask: Perplexity's open Decider in place of jev on 92 company questions
hedwigai ask answers a question by looking it up and deciding as it goes: which lookups to run, which figure measures the amount asked about, and whether the answer is yes or no. Those decisions are made by jev, a decision model hedwigai serves. We ran the same 92 yes-or-no company questions with every one of them made instead by Decider, Perplexity's open decision model, on two NVIDIA RTX 3090s of our own. Every answer was checked against SEC records.
Decider gets as many right as jev, 89 of 92, with one more wrong. It is about 9 times slower on those two cards, and it costs nothing per decision.
| ask with jev | ask with Decider | |
|---|---|---|
| Right | 89 | 89 |
| Wrong | 1 | 2 |
| Unsure | 2 | 1 |
| Right, of those answered | 99% | 98% |
| Time, median per question | 2.1 s | 19.1 s |
| Decisions per question | 81 in 2.2 calls | 81 in 2.2 calls |
| Where the decisions run | hedwigai's jev service | two NVIDIA RTX 3090s |
| Charge per decision | billed per call | none; the GPUs are ours |
Question by question
The two agree on almost every kind. Decider answers all 15 large biotechs where jev is unsure of 2. It answers yes to one company that is not a biotech, because the lookup it chose never asked whether it was one, and is unsure of one whose question names the wrong stock exchange.
Where the time goes
A question takes about 81 decisions, sent in 2.2 calls on average; most of them come at once, when ask plans which lookups to run. jev answers a call like that in well under a second. Decider is a 27-billion-parameter model split across two consumer cards and asked 16 decisions at a time, so the planning call alone takes most of its 19.1 s. That is the cost of these two cards, not of the model: a single larger card, or more of them, cuts it.
One sentence decided 22 answers
Every question here asks about a market cap or a public float, and the only lookup that returns either is the one for shares outstanding and the public float. To check an amount, ask first decides which lookup measures it, or the closest stand-in. Whatever it picks is the only number the comparison gets; if it picks none, the amount cannot be checked and the answer stays unsure.
We ran both systems with that lookup described two ways. Said as "NOT a market cap", Decider chose "none of these" on 59 of 92 questions, narrowly each time (about 0.55 against 0.39), and left 23 answers unsure. Said as the closest stand-in for a market cap, it chose the lookup at 0.96. jev chose it either way, and its answers did not change. Decider reads a description literally; the descriptions ask decides over now say what each lookup is good for, not only what it is not.
- Not a market cap
- shares outstanding and the public float, dated and citable — NOT a market cap, which needs a share price nothing here has
- The stand-in (used now)
- shares outstanding and the public float, dated and citable — the public float is the closest stand-in for a market cap here; a market cap itself needs a share price nothing here has
Fine print
- The questions, their kinds and the SEC ground truth are the 92 of jev, Vela 2.0 and ask on 92 yes/no questions. "Biotech" means SEC industry codes 2833 to 2836 or 8731.
- Both arms ran the same build of ask, on the same day (10 October 2026), with the same lookups; only the service answering the decisions differed. Each decision was asked as jev is asked, and Decider's answer given back in jev's form, read with the same thresholds.
- Decider is pplx-decider-v1-27b, Perplexity's open weights, run in 8-bit weights on two NVIDIA RTX 3090s (24 GB each), 16 decisions at a time.
- Times are medians of the whole question, lookups included, which are the same for both; the difference between the two is the decisions.
- One kind of question, one account, 92 questions, one run each way.
More head-to-head evals
| Measure | jev | ask |
|---|---|---|
| Our 370 clinical decisions, right | 94.6% | 93.0% |
| PubMedQA, right (best EdgeEvals scored: 62.2%) | 77.0% | 74.4% |
| Measure | ask | Claude Opus 5.5 |
|---|---|---|
| Questions right, of 95 | 59 | 63 |
| Small companies right, of 20 | 13 | 9 |
| Measure | realestate-us | Best of four assistants |
|---|---|---|
| Floor-plan rents right, of 26 | 26 | 23 |
| Amenities claimed the listing doesn't state | 12 | 129 |
| Measure | ask | jev |
|---|---|---|
| Right, of the questions answered | 98% | 79% |
| Wrong answers, of 92 | 2 | 19 |
| Measure | Trained localiser | Best of four frontier models |
|---|---|---|
| Mean overlap with the defect (IoU, 0 to 1) | 0.421 | 0.109 |