
A vendor demonstration searches for an exact title and finds it. So does every other vendor. The dataset that separates them is the one built from your own catalogue.
The short answer: no public dataset predicts how an enterprise search tool will behave on your content, because the hard part of your search is your vocabulary and your users' mistakes. Generate the dataset yourself: sample a few hundred items you know exist, produce the ways people actually ask for them in 20 variation classes, add a nonsense set that must return nothing, and add your real failed queries. Score every candidate per class, never only overall.
This is the method ITMTB uses to evaluate its own search engine on every change and to compare it with whatever a customer had before. It pairs with how to measure the success of an enterprise search tool, which covers what to measure once a tool is live.
Retrieval research has good public datasets: MS MARCO, the BEIR collection, the TREC tracks. They are the right way to compare retrieval models in general and the wrong way to choose an enterprise search tool, for three reasons.
Take an item you know exists. Damage the query the way a real person would, keeping the intent intact. Ask whether the item still comes back in the top five.
That is the whole method. The correct answer is known by construction, so no human relevance judgement is needed, and the set can be regenerated at any size on any catalogue.
These are the classes the harness generates from each sampled title. Some apply only to titles with enough words, a plural or a hyphen, so their counts will be lower than the sample.
| Class | What it does to the title | What it tests |
|---|---|---|
| exact_title | Nothing | Baseline; should be close to 100% |
| subject | Keeps the subject, drops the catalogue's convention words | Convention handling |
| lowercase | All lower case | Case handling |
| typo_delete | Deletes one letter | Fuzzy matching |
| typo_insert | Inserts one letter | Fuzzy matching |
| typo_substitute | Replaces one letter with a keyboard neighbour | Realistic typos |
| typo_transpose | Swaps two adjacent letters | Realistic typos |
| typo_plural_noise | Adds or removes a trailing s | Stemming |
| typo_phonetic | Spells a word as it sounds | Phonetic correction |
| typo_two_words | Misspells two words at once | Correction under load |
| order_reverse | Reverses word order | Order independence |
| order_rotate | Moves the first word to the end | Order independence |
| order_shuffle | Random order | Order independence |
| drop_word | Removes one content word | Partial recall |
| nl_phrasing | Wraps the title in natural language | Mandatory-word handling |
| punct_strip | Removes punctuation | Tokenisation |
| hyphen_join | Joins or splits hyphenated terms | Tokenisation |
| plural_flip | Singular to plural or back | Stemming |
| combo_order_typo | Reorders and misspells together | Compound damage |
| abbrev | Abbreviates a multi-word term | Abbreviation knowledge |
The per-class table is what finds the work. On the catalogue where this harness runs in production, the overall top-5 figure is 98.8%, and the class that drove the most engineering was abbreviations, which started around 90%. An overall figure would have hidden it.
Generated classes test what you anticipated. Once a search is live, people complain about things you did not. Each complaint should become a new class, generated the same way and scored outside the overall figure so it cannot be averaged away.
Classes that have been added this way include: a query made entirely of the catalogue's convention words plus a subject; a single wrong vowel in a technical term; one word more than the title contains; very short subjects that must rank the exact item first; hyphenated compounds that must not match either half on its own; and a geography or year qualifier the catalogue does not use in its titles. Real traffic contained few of those queries, which is why no generated class had caught them and why a dataset needs a channel for complaints.
A search that returns something for everything has not solved retrieval; it has hidden the zero. Keep a set of queries that must return nothing: invented words, and phrases that combine two unrelated subjects. Score the number of results returned. Good is zero.
This set matters most once a meaning-based component is in the path, because embeddings will find the nearest neighbour of nonsense. A reasonable rule is that the meaning-based channel may only answer when the lexical engine found nothing, and only above a confidence floor; the nonsense set is what calibrates the floor.
The third part of the dataset is not generated. Every query from your logs that returned nothing, and every query a user complained about, goes in with the item it should have found. It runs on every change, versioned, and grows with every incident. A fix for one complaint that breaks another shows up here first.
If you are comparing vendors and have no logs yet, collect them first. A week of real queries says more than any amount of generated data about which classes your users actually hit.
| Score | Definition | Rule |
|---|---|---|
| top-5, top-10 per class | Share of cases where the known item appears in the first 5 / 10 results | Report per class. A class dropping while overall holds is a regression. |
| top-1 for exact classes | For exact titles and exact subjects, the item must be first | Below about 95% means ranking, not recall, is the problem |
| p50 latency per class | Median engine time per class | Record on the same hardware for every candidate; watch for classes that cost ten times others |
| Nonsense leaks | Results returned for queries that should return none | Zero |
| Replay pass rate | Real failed queries that now find their item | 100%, and the corpus must grow |
Two things to know before trusting the numbers:
| Step | Output |
|---|---|
| Export item identifiers and titles, plus the one or two fields users search by | A CSV of your corpus |
| Sample a few hundred items at random with a fixed seed | The known-item list |
| Generate the 20 classes per item with a script; the typo classes need only a keyboard-neighbour map and a vowel table | Query, expected item and class, several thousand rows |
| Write a few dozen nonsense queries | The guardrail set |
| Take recent zero-result queries from your logs and find the intended item by hand for the most frequent | The replay set |
| Run each candidate; record the rank of the expected item and the latency per query | Raw results |
| Aggregate per class: top-5, top-10, p50; count leaks; count replay passes | The comparison table |
Once the engine is live, the same dataset becomes the gate on every later change, and the live measures in how to measure enterprise search success take over from there. Where the engine also serves AI agents as a tool, run the dataset through the tool interface as well; an agent should see exactly what the search box sees.
ITMTB generates this dataset from a prospective customer's catalogue and runs it against their current search before scoping any build, so the decision rests on a per-class table rather than a demonstration; see the enterprise search page. Contact us.
Public retrieval benchmarks exist, but they score an engine on general documents and question-style queries. They do not contain your items, vocabulary or your users' mistakes, so they cannot predict behaviour on your catalogue.
A few hundred items generating several thousand queries across 20 classes detects a one-point change in a class on a catalogue of tens of thousands of items. Use two seeds.
Known-item queries in realistic variation classes, a nonsense set that must return nothing, and your real failed queries with the item each should have found.
Same corpus, dataset, depth, hardware budget and day. Score top-5, top-10 and latency per class, leaks on nonsense, passes on replay. Compare per class.
Because users rarely type them, and every engine passes. The realistic classes are where engines differ.
Yes, with the nonsense guardrail weighted heavily, because meaning-based matching tends to answer everything.
The class list and scores are from the evaluation harness of ITMTB's enterprise search for The Business Research Company, described in the case study.
ITMTB generates the evaluation dataset from your catalogue, adds your real failed queries, and runs it against the candidates on the same corpus and budget. The output is a per-class table.