A Realistic Dataset for Comparing Enterprise Search Solutions

Public benchmarks measure an engine on somebody else's documents. The dataset that predicts your results is generated from your own items, your users' mistakes and your failed searches. The method, the query classes and the scoring rules.

A Realistic Dataset for Comparing Enterprise Search Solutions

A vendor demonstration searches for an exact title and finds it. So does every other vendor. The dataset that separates them is the one built from your own catalogue.

The short answer: no public dataset predicts how an enterprise search tool will behave on your content, because the hard part of your search is your vocabulary and your users' mistakes. Generate the dataset yourself: sample a few hundred items you know exist, produce the ways people actually ask for them in 20 variation classes, add a nonsense set that must return nothing, and add your real failed queries. Score every candidate per class, never only overall.

This is the method ITMTB uses to evaluate its own search engine on every change and to compare it with whatever a customer had before. It pairs with how to measure the success of an enterprise search tool, which covers what to measure once a tool is live.

Why public benchmarks do not answer the question

Retrieval research has good public datasets: MS MARCO, the BEIR collection, the TREC tracks. They are the right way to compare retrieval models in general and the wrong way to choose an enterprise search tool, for three reasons.

  • They do not contain your vocabulary. The failures that matter in a specialist catalogue are specialist: a general dictionary correcting a valid technical term, a two-letter abbreviation missing from the index, a hyphenated product name split into two halves.
  • They do not contain your users' mistakes. People searching a particular catalogue misspell its words, drop its convention words and reorder its titles in patterns visible only in its logs.
  • They reward a different task. Most public benchmarks score ranked relevance over long documents for question-style queries. Most enterprise search is known-item retrieval: the user wants one specific thing and cannot name it exactly.

The principle: known items, realistic damage

Take an item you know exists. Damage the query the way a real person would, keeping the intent intact. Ask whether the item still comes back in the top five.

That is the whole method. The correct answer is known by construction, so no human relevance judgement is needed, and the set can be regenerated at any size on any catalogue.

The 20 variation classes

These are the classes the harness generates from each sampled title. Some apply only to titles with enough words, a plural or a hyphen, so their counts will be lower than the sample.

ClassWhat it does to the titleWhat it tests
exact_titleNothingBaseline; should be close to 100%
subjectKeeps the subject, drops the catalogue's convention wordsConvention handling
lowercaseAll lower caseCase handling
typo_deleteDeletes one letterFuzzy matching
typo_insertInserts one letterFuzzy matching
typo_substituteReplaces one letter with a keyboard neighbourRealistic typos
typo_transposeSwaps two adjacent lettersRealistic typos
typo_plural_noiseAdds or removes a trailing sStemming
typo_phoneticSpells a word as it soundsPhonetic correction
typo_two_wordsMisspells two words at onceCorrection under load
order_reverseReverses word orderOrder independence
order_rotateMoves the first word to the endOrder independence
order_shuffleRandom orderOrder independence
drop_wordRemoves one content wordPartial recall
nl_phrasingWraps the title in natural languageMandatory-word handling
punct_stripRemoves punctuationTokenisation
hyphen_joinJoins or splits hyphenated termsTokenisation
plural_flipSingular to plural or backStemming
combo_order_typoReorders and misspells togetherCompound damage
abbrevAbbreviates a multi-word termAbbreviation knowledge

The per-class table is what finds the work. On the catalogue where this harness runs in production, the overall top-5 figure is 98.8%, and the class that drove the most engineering was abbreviations, which started around 90%. An overall figure would have hidden it.

Classes that come from live complaints

Generated classes test what you anticipated. Once a search is live, people complain about things you did not. Each complaint should become a new class, generated the same way and scored outside the overall figure so it cannot be averaged away.

Classes that have been added this way include: a query made entirely of the catalogue's convention words plus a subject; a single wrong vowel in a technical term; one word more than the title contains; very short subjects that must rank the exact item first; hyphenated compounds that must not match either half on its own; and a geography or year qualifier the catalogue does not use in its titles. Real traffic contained few of those queries, which is why no generated class had caught them and why a dataset needs a channel for complaints.

The nonsense guardrail

A search that returns something for everything has not solved retrieval; it has hidden the zero. Keep a set of queries that must return nothing: invented words, and phrases that combine two unrelated subjects. Score the number of results returned. Good is zero.

This set matters most once a meaning-based component is in the path, because embeddings will find the nearest neighbour of nonsense. A reasonable rule is that the meaning-based channel may only answer when the lexical engine found nothing, and only above a confidence floor; the nonsense set is what calibrates the floor.

Real failed searches

The third part of the dataset is not generated. Every query from your logs that returned nothing, and every query a user complained about, goes in with the item it should have found. It runs on every change, versioned, and grows with every incident. A fix for one complaint that breaks another shows up here first.

If you are comparing vendors and have no logs yet, collect them first. A week of real queries says more than any amount of generated data about which classes your users actually hit.

How to score

ScoreDefinitionRule
top-5, top-10 per classShare of cases where the known item appears in the first 5 / 10 resultsReport per class. A class dropping while overall holds is a regression.
top-1 for exact classesFor exact titles and exact subjects, the item must be firstBelow about 95% means ranking, not recall, is the problem
p50 latency per classMedian engine time per classRecord on the same hardware for every candidate; watch for classes that cost ten times others
Nonsense leaksResults returned for queries that should return noneZero
Replay pass rateReal failed queries that now find their item100%, and the corpus must grow

Two things to know before trusting the numbers:

  • Ties flicker. Items with identical relevance scores can come back in a different order after an index rebuild, and a case sitting on the top-5 boundary flips. Record the index build with every score, sort ties deterministically if you need stability to a decimal, and treat differences below the noise as noise.
  • Overall hides the class you care about. A change that lifts long queries can cost the drop-a-word class and move the overall figure by less than the flicker. Only the per-class table shows it.

Comparing vendors fairly

  1. Same corpus. Give every candidate the identical export of your items, same fields, same count.
  2. Same dataset. Generate it once, from one seed, and hand the same file to each. Keep a second seed back to check the result is not an artefact of the first.
  3. Same depth and budget. Ten results, same hardware class, same timeout. A tool that is excellent at two seconds and poor at 300 milliseconds is a different tool.
  4. Per class, side by side. Put the candidates' per-class tables next to each other and weight the classes your logs say matter.
  5. Nonsense and replay last. A candidate that leaks on nonsense or fails your real queries is out, whatever its overall score.
  6. Ask how failures are counted. The tool you choose will fail sometimes. How a failed, degraded or timed-out search is counted and shown decides whether you will know.

Building the dataset

StepOutput
Export item identifiers and titles, plus the one or two fields users search byA CSV of your corpus
Sample a few hundred items at random with a fixed seedThe known-item list
Generate the 20 classes per item with a script; the typo classes need only a keyboard-neighbour map and a vowel tableQuery, expected item and class, several thousand rows
Write a few dozen nonsense queriesThe guardrail set
Take recent zero-result queries from your logs and find the intended item by hand for the most frequentThe replay set
Run each candidate; record the rank of the expected item and the latency per queryRaw results
Aggregate per class: top-5, top-10, p50; count leaks; count replay passesThe comparison table

Once the engine is live, the same dataset becomes the gate on every later change, and the live measures in how to measure enterprise search success take over from there. Where the engine also serves AI agents as a tool, run the dataset through the tool interface as well; an agent should see exactly what the search box sees.

ITMTB generates this dataset from a prospective customer's catalogue and runs it against their current search before scoping any build, so the decision rests on a per-class table rather than a demonstration; see the enterprise search page. Contact us.

Frequently asked questions

Is there a public dataset for comparing enterprise search solutions?

Public retrieval benchmarks exist, but they score an engine on general documents and question-style queries. They do not contain your items, vocabulary or your users' mistakes, so they cannot predict behaviour on your catalogue.

How big does it need to be?

A few hundred items generating several thousand queries across 20 classes detects a one-point change in a class on a catalogue of tens of thousands of items. Use two seeds.

What should it contain?

Known-item queries in realistic variation classes, a nonsense set that must return nothing, and your real failed queries with the item each should have found.

How do you compare vendors fairly?

Same corpus, dataset, depth, hardware budget and day. Score top-5, top-10 and latency per class, leaks on nonsense, passes on replay. Compare per class.

Why not test with exact titles?

Because users rarely type them, and every engine passes. The realistic classes are where engines differ.

Should semantic or AI search use the same dataset?

Yes, with the nonsense guardrail weighted heavily, because meaning-based matching tends to answer everything.

Key takeaways

  • No public dataset predicts your enterprise search outcome. Generate one from your own items, in realistic variation classes, plus nonsense and your real failures.
  • A few hundred items, several thousand queries, two seeds and 20 classes is enough to see one-point changes.
  • Score per class. A class dropping while overall holds is a regression.
  • Zero leaks on nonsense. A search that answers everything has hidden the zero, not solved it.
  • Add every real complaint as a new class.
  • Compare vendors on the same corpus, dataset, depth and budget, per class, and ask how they count their own failures.

Source note

The class list and scores are from the evaluation harness of ITMTB's enterprise search for The Business Research Company, described in the case study.


Comparing search tools on your own catalogue?

ITMTB generates the evaluation dataset from your catalogue, adds your real failed queries, and runs it against the candidates on the same corpus and budget. The output is a per-class table.

Explore More Insights

How to Measure the Success of an Enterprise Search Tool: 12 Measures From a Live Deployment

How to Measure the Success of an Enterprise Search Tool: 12 Measures From a Live Deployment

Read More
Enterprise Search Case Study: How ITMTB Took The Business Research Company's Zero-Result Searches From 53% to 5%

Enterprise Search Case Study: How ITMTB Took The Business Research Company's Zero-Result Searches From 53% to 5%

Read More
Improving Enterprise Search for People and AI Agents

Improving Enterprise Search for People and AI Agents

Read More
Why Enterprise Search Is Important: The Cost of Content Nobody Can Find

Why Enterprise Search Is Important: The Cost of Content Nobody Can Find

Read More