Workbooks for everyday work.
Lead reports, due diligence documents and spreadsheets.
6 minutes later: qualified B2B sales leads across five product lines.
And every source it cited was really where it said it was. Not most of them. Every one.
See the run →Works inside Claude, Codex and Cursor.
Your workbooks are there in the apps you already work in: ask a question, start a report, or update a spreadsheet right where you're at.
- Claude Code
- Claude
- Codex
- Cursor
- Hermes
- Any MCP app
Written from primary sources and your own files.
A document is only as good as what it was written from. So a workbook goes to the primary sources that matter to you: SEC filings, FDA labels, clinical trial registries, customs records, procurement awards.
And to the files you already keep. Connect Google Drive, Dropbox, OneDrive or SharePoint, and it works from what is there.
If a source you need is missing, we add that data hose. That is our enterprise work: talk to us.
Every line cites the page or row it came from.
How it works.
- 01
Describe it. One sentence saying what the document is.
- 02
It does the work. It reads the primary sources and your own files, and cites each one.
- 03
Several AI models share it. Each takes the part it does best: gathering, writing, checking.
- 04
It updates itself. When a source changes, the workbook changes with it and tells you what moved.
- 05
You change it with a sentence. Ask in plain English and it rewrites the part you meant.
Connect your AI agent with MCP.
Connect the hedwigai MCP and your agent can ask, list, start and change workbooks. You speak in sentences; it makes the calls.
A short answer from your workbooks and programs, with the source it rests on. Nothing to wait for.
Is Moderna a biotech with a public float above $5B?
Your agent
ask · Is Moderna NASDAQ:MRNA a biotech company with public float above $5B?
{
"answer": "yes",
"claims": [{ "figure": "public float", "value": "$9.6B",
"as_of": "2026-06-28" }],
"sources": ["MRNA 10-K, cover page",
"biotech-diligence-7fq2 · §Moderna"]
}Yes. Moderna's public float was $9.6B as of June 28, per the cover of its 10-K.
Lead reports, due diligence and spreadsheets.
A workbook is one document with one job. You don't maintain it; it arrives researched and finished, and it keeps itself up to date.
Lead reports
Who to sell to, and why now. Customs records, procurement awards and project wins, qualified against your own criteria, with an outreach email drafted for each.
US importers of heavy engineering equipment.
Diligence documents
The pack on a company, a target or a supplier. Filings, labels, trial results and investor decks, gathered and checked, with every claim linked to its source.
Diligence on eleven biotech companies.
Spreadsheets
Trackers, budgets and models. Change an assumption in plain English; the numbers move and it tells you which ones did.
The client list nobody has updated since March.
You can check every line.
Sources on every claim
Each line traces to the filing, page or row it came from.
Every update says what changed
When a source moves, you see which lines changed and which source moved them.
Shared with your team
Invite the people you work with, and they can read the workbook and ask for changes too.
You are asked only when it matters
It stops for a decision or an approval, and otherwise gets on with the work.
Tested on real work. Every score published.
Each test is a real job, scored against a checklist written before the run started. We publish every result, including what it missed.
Time machine: leak-free RL environments, graded by what actually happened
A feature: pick a date, and every source answers as it would have on that day.
| Measure | RL player | A coin |
|---|---|---|
| Score, clock before the answer is measured | 0.747 | 0.750 |
| Score, clock after it (a leak) | 0.762 | 0.750 |
Self-hosted decisions for ask: Perplexity's open Decider in place of jev on 92 company questions
The decisions inside hedwigai ask made by Perplexity's open Decider on two NVIDIA RTX 3090s, against jev.
| Measure | ask with jev | ask with Decider |
|---|---|---|
| Questions right, of 92 | 89 | 89 |
| Wrong answers | 1 | 2 |
Biotech due-diligence document gathering: coverage across 11 companies
Eleven companies of diligence documents, gathered and scored.
| Companies covered | 11 |
|---|---|
| Mean coverage, 0 to 1 | 0.971 |
US importer search for heavy engineering equipment: coverage across 5 product lines
Five product lines of importers found, qualified and gated.
| Product lines covered | 5 |
|---|---|
| Mean coverage, 0 to 1 | 0.956 |
Aircraft records completeness: an evaluation method (no models scored yet)
Whether a records package is complete — the method, before the scores.
| Methods described | 8 |
|---|---|
| Kinds of defect | 9 |
Clinical decisions: jev and ask on System One, PubMedQA, SciFact and a 370-decision set
EdgeEvals' System One board, PubMedQA, SciFact, and a clinical set built to System One's rules.
| Measure | jev | ask |
|---|---|---|
| Our 370 clinical decisions, right | 94.6% | 93.0% |
| PubMedQA, right (best EdgeEvals scored: 62.2%) | 77.0% | 74.4% |
Executive lookup for US-listed companies: Claude Opus 5.5 and ask on 95 questions
Who runs a public company, and how to email them: a frontier model from memory against hedwigai ask.
| Measure | ask | Claude Opus 5.5 |
|---|---|---|
| Questions right, of 95 | 59 | 63 |
| Small companies right, of 20 | 13 | 9 |
Rental comps: floor-plan rents and amenities on 14 listings, five assistants compared
Rent per floor plan and amenities for US rentals: hedwigai realestate-us against Claude, GPT, Gemini and Perplexity with web search.
| Measure | realestate-us | Best of four assistants |
|---|---|---|
| Floor-plan rents right, of 26 | 26 | 23 |
| Amenities claimed the listing doesn't state | 12 | 129 |
Company screening from SEC filings: jev, Vela 2.0 and ask on 92 yes/no questions
Yes-or-no questions about public companies, answered by two decision models alone and by hedwigai ask.
| Measure | ask | jev |
|---|---|---|
| Right, of the questions answered | 98% | 79% |
| Wrong answers, of 92 | 2 | 19 |
A same-model decider above an agent's writer: one task, two configurations
The same document and the same instruction, run with and without a decider above the writer.
| Measure | With a decider | Without |
|---|---|---|
| Wall clock | 23m 36s | 3m 46s |
| Model calls | 104 | 34 |
Defect localization in industrial images: four frontier models and a trained localizer
Four frontier models asked to put a box around what they just described.
| Measure | Trained localiser | Best of four frontier models |
|---|---|---|
| Mean overlap with the defect (IoU, 0 to 1) | 0.421 | 0.109 |
Company lists in ask: a small model trained to put the right companies first, tested on 33 new questions
Ask for a list of companies and get ten, each with its source. How the ten are chosen, and how much better it got.
| Measure | Trained for lists | General-purpose model |
|---|---|---|
| Correct, of the companies in the top 10s | 62% | 55% |
| Correct companies in the top 10s | 142 | 117 |
