jev against yantra on 92 questions
Some questions want a word, not a document. Is this company a biotech. Is its float above ten billion dollars. Does the ticker even belong to the name. Starting a research run for each of those is the wrong tool, so we built a short-answer path and measured it against the obvious alternative: asking a model directly.
Of 92 questions, yantra answered 66 and got 65 of them right, 98%. jev, asked directly, answered all 92 and got 73 right, 79%. yantra was wrong 1 time; jev was wrong 19 times. yantra also said it was unsure 26 times, and took about 11 times as long.
The two systems
jev is a decision model: it takes a question and a set of options and answers with a probability for each. Here it was given the question alone and three options: yes, no, or the premise is false (no such company, or the name and ticker do not match).
yantra is hedwigAI's own system for the same question. Small models on our CPU inference service read the question: GLiNER finds the company name and ticker in about twenty milliseconds. jev then decides only what the small models cannot: which of the account's programs to run, and whether the evidence they return answers the question. The programs are the same ones our research runs use, here reading SEC filings. The account's own workbooks are searched alongside, and every answer cites the program output or workbook section it rests on.
Two rules matter for the results. Numbers are never compared by a model, only by code. And a name and ticker that lead to two different records, or to none, is reported as a false premise, not as a no.
The questions
Every question has the same shape, written several ways: is this company, with this exchange and ticker, a biotech company with a market cap or public float above some amount. The answers were computed from SEC records before either system ran. Biotech means the company files under SEC industry codes 2833 to 2836 or 8731. Size means the latest public float the company reported.
Float is never more than market cap, so a question worded as market cap only has a threshold where the answer holds either way. A threshold close to the float is only ever worded as public float. Ten of the companies and tickers are invented, and ten pair one company's ticker with another's name.
Results
Right · wrong · unsure
| Kind of question | n | jev | yantra, first run | yantra, second run |
|---|---|---|---|---|
| Invented companya name and ticker that do not exist | 10 | 9 · 1 · 0 | 10 · 0 · 0 | 10 · 0 · 0 |
| Name and ticker disagreeone company's ticker under another's name | 10 | 10 · 0 · 0 | 10 · 0 · 0 | 10 · 0 · 0 |
| Wrong exchangea real company on the other exchange | 9 | 7 · 2 · 0 | 5 · 2 · 2 | 5 · 0 · 4 |
| Biotech, clearly abovethe threshold well under the float | 15 | 8 · 7 · 0 | 12 · 2 · 1 | 11 · 1 · 3 |
| Biotech, clearly belowthe threshold over three times the float | 15 | 14 · 1 · 0 | 14 · 1 · 0 | 10 · 0 · 5 |
| Biotech, near the linethe threshold within 6% of the float | 10 | 3 · 7 · 0 | 10 · 0 · 0 | 8 · 0 · 2 |
| Not biotechany size; the answer is no | 23 | 22 · 1 · 0 | 12 · 7 · 4 | 11 · 0 · 12 |
| All | 92 | 73 · 19 · 0 | 73 · 12 · 7 | 65 · 1 · 26 |
| Right, of those answered | 79% | 86% | 98% |
jev is good at what can be judged from the words alone. It knew the invented companies were invented and caught every mismatched ticker. It is weak wherever the answer depends on a number it cannot see: it got 3 of 10 right near the line and 8 of 15 biotechs clearly above it.
yantra's first run won exactly there, 10 of 10 near the line, because it read the figure from the filing and compared it in code. It lost on companies that are not biotech: 12 of 23, against jev's 22.
What the first run got wrong
Ten companies that are not biotech were called biotech. jev was shown the question with its numeric condition already settled beside it ("float above the amount: yes"), and leaned toward yes on the rest. One was called biotech at 0.86. The fix was to stop showing it: jev now judges the question with the amount clause cut out, and is not asked at all when nothing is left to judge.
Six real companies were reported as false premises. Most were lookups that ran past their four second budget and were counted as having found nothing. A lookup that did not answer now says nothing about whether a company exists. A name that misses where its own ticker finds the same company is now a lookup's miss, not a false premise.
After both fixes, wrong answers fell from 12 to 1. The last one is a question where the name was read as "Nasdaq: ERAS biotech".
Why yantra says unsure
The second run said unsure 26 times, and in 21 of them one cause was present. The program that reads a company's float from its filings took longer than its four seconds the first time it saw that company, so the amount could not be settled and the answer stayed open. The budget was set for a warm lookup. That is a setting, not a limit of the method.
The rest are jev declining to call a beverage or a bank a biotech with confidence, once the company's name is taken out of what it reads. We take the name out so that one decision about an industry serves every company in it. Here that cost confidence on the easy cases.
An unsure is not a free pass. It is a question someone still has to answer. But it is the honest answer when the evidence did not come back, and it is never a confident wrong one.
Time
jev answered in 0.5 s at the median. yantra took 5.3 s, 4.8 s of it on the server. The small models and jev's decisions together come to about 0.7 s; most of the rest is the programs reading SEC filings, often for the first time. Once a company has been looked up, the same question takes about 1.6 seconds on the server, and a question answered from the account's own workbooks takes under half a second.
What this does not show
One question shape, one account and 92 questions. Every question names a public company that files with the SEC, which is where yantra's programs are strong; a question its programs cannot reach would leave it unsure, and jev would still answer.
"Biotech" here is an SEC industry code, which counts large drug makers in and leaves out companies people might call biotech. Both systems were scored against the same definition, which neither was told.
Both arms were timed from one laptop over the public internet, one question at a time. yantra has not yet beaten jev on the measure we set for switching to it: clearly better on both accuracy and speed. On accuracy, of what it answers, it is. On speed it is not.
Measured on 6 and 7 October 2026, 92 questions, labels computed from SEC records. Every reply, with its timings and sources, is kept with the question set.