hedwigaihedwigaihedwigai
Blog

Company lists in ask: a small model trained to put the right companies first, tested on 33 new questions

Ask hedwigai ask for a list of companies, like “Which food distributors supply restaurants in Goa?”, and it searches the web, reads the pages it finds, and answers with ten companies, each with the sentence and the page it came from.

Which ten, and in what order, is decided by a small model. We trained one for this job. On 33 questions it had never seen, 62% of the companies it put in the top ten were correct, up from 55%.

What you get

ask is a tool on hedwigai's MCP server, so it works inside the AI assistant you already use. You type a question in plain words. The answer is a list you can check, because every company comes with the sentence that names it and a link to the page.

The first three companies of a list from a live run:

CompanyWhat its page says
Rodaaji CompanyImporter, brand representative and distributor
Nuvo Mart…a leading distributor for high quality food items…
Ashokasha…works across retail, distribution and professional food…

“Which HoReCa foodservice distributors supply restaurants and hotels in Goa, India?”, October 10, 2026. A list takes about half a minute.

How a list is made

  1. Search. Your question is turned into several web searches.
  2. Read. One page from each site that comes back is read, up to fourteen pages.
  3. Find the names. Every organization the pages mention is picked out, and different spellings of one company (“TJUK”, “TJUK Trade Networks Pvt. Ltd.”) are merged.
  4. Check what each name is. People, places and events are removed, and so are investors and newspapers unless that is what you asked for.
  5. Put them in order. Each company is scored on how well the sentence about it answers your question. The best ten are your answer.

This post is about step 5. Steps 3 to 5 each run on a small model on our own servers, not on a large language model.

Teaching the model

Step 5 used a general-purpose ranking model, built to match search queries to web passages. It had never seen a question like “which distributors operate in Mumbai”. To teach it, we needed examples marked right or wrong.

We ran 132 list questions and kept every company each one considered: 2,213 examples, each a company name and the sentence about it. Two large AI models, jev and gpt-oss-120b, marked each one. We kept the 1,613 examples they agreed on and trained the small model on those.

A mistake both made

One distributor's website says “We supply 100+ premium brands including McCain, Rich's, Pillsbury, Barry Callebaut.” Asked whether McCain is a food distributor in Mumbai, both AI models said yes. Because they agreed, the mistake went into the training examples, and into the examples we were testing against.

So we changed the question. Instead of “is this what the user asked for?”, each model was asked what the name is in that sentence: the kind of company asked for, a brand that company sells, a customer, a person, a place, and so on. Asked that way, both called McCain a brand, and the distributor whose website it is stayed a distributor. The model now live was trained on these answers.

Results

We wrote 33 questions the model had not been trained on (frozen food distributors in Pune, hotel chains in Goa, wine distributors in Singapore and so on), ran each once, and had the same two AI models mark every company found, 1,225 in all. Then we put each question's companies in order both ways, before and after, and counted what landed in the top ten.

Correct in the top 10sWrong in the top 10sShare correctLists better / worse
Before: a general-purpose ranking model1179755%—
After: the model trained for lists (live now)1428662%18 / 5

Across the 33 lists, the trained model put 25 more correct companies in the top tens and 11 fewer wrong ones. It gave a better list for 18 questions and a worse one for 5, each of those by one company. Companies the two AI models disagreed about count as neither correct nor wrong.

On the Mumbai question, the old model's top ten still had McCain, Pillsbury and Barry Callebaut in it. The new one's has none of them.

What can still go wrong

  • “Correct” here means two AI models agreed, not that a person checked. They can share a mistake, as they did with brands, and a shared mistake does not show up in these numbers.
  • Every run reads different pages. Ask the same question twice and the list can change. The before and after comparison used the same companies for both, so it is fair, but your list will not match ours.
  • Some wrong names get through. A company that bought another, named in news of the deal, can look like the kind you asked for, and a town written on its own line in an address can look like a company.
  • 33 questions is a small test. The gain was clear, better on 18 and worse on 5, but the exact percentages would move with another 33.

Since this test, each company is scored on its best sentence instead of its first, and investors and newspapers are removed more reliably. Those changes are live, and not in the figures above.

Tested October 2026 on 33 questions and 1,225 companies. The trained model is MiniLM (22 million parameters), fine-tuned from cross-encoder/ms-marco-MiniLM-L6-v2. The two AI models that marked the examples are jev and gpt-oss-120b.