Blog
Petal+Eon hosts floral workshops where people press real preserved flowers into a BloomFrame. BloomFrameAI turns a text prompt or an uploaded photo into custom artwork for it — within the frame's margins, around the flower area, and ready for the die cutter — so every frame can be personal without being made by hand.
- Workshops hosted
- 1,000+
- Frame constraints met
- 5
- Stages per design
- 3
- Customer
- Petal+Eon
How completely an AI agent gathers biotech due diligence documents — SEC filings, FDA drug labels, ClinicalTrials.gov entries, investor presentations and the published literature — measured across 11 companies.
- Companies
- 11
- Mean coverage
- 0.971
- Wall clock
- 3h 11m
- Tokens
- 55.9M in · 241k out
How completely an AI agent finds and qualifies US importers of heavy engineering equipment — customs bill-of-lading records, utility and public-power procurement, EPC project awards, and the certification and duty gates that decide whether a lead can convert — measured across five product lines.
- Product lines
- 5
- Mean coverage
- 0.956
- Wall clock
- 35m 35s
- Tokens
- 6.67M in · 76k out
Whether a records package is complete: life-limited part trace, directive compliance, repair approvals, release certification and the lease's own return conditions. The method and its scoring rules, with a worked example of what a scored run reports. No model has been measured on it yet.
- Methods
- 8
- Defect taxonomy
- 9 kinds
- Seeded per package
- 82
- Corpora
- synthetic · NDA gold set
We ran an agent task twice with one variable: whether a second system sat above the model choosing its tool calls. Same document, same instruction, same writer. The expensive arm produced an identical document, mistakes included, for ten times the prompt tokens. Why that happened and what we changed.
- Wall clock
- 23m 36s vs 3m 46s
- Model calls
- 104 vs 34
- Prompt tokens
- 6.62M vs 620k
- Output difference
- none
For anyone choosing how to read industrial images: a frontier model can describe a weld, a casting or a part, but it cannot put a box around what it just described. Four models, one verdict and one box per image, against a rule that never looks and a small model trained for the job.
- Models
- 4
- Best mean IoU
- 0.109
- Trained localiser
- 0.421
- Median IoU
- 0.000