
Someone asks whether the new search is better. Everyone has an opinion. Nobody has a number that would survive a second question.
The short answer: ITMTB measures enterprise search on three layers. Offline: a known-item test where realistic variations of items you know exist must return that item in the top five, tracked per variation class. Live: zero-result rate deduplicated by intent, fallback and degraded counters, latency against a budget, and clicks attributed to the search that produced them. Business: catalogue reach, the share of items that search ever surfaces, and the demand visible in honest zero results. Most of the difficulty is in the counting.
This article sets out the twelve measures ITMTB uses for an enterprise search deployment: what each one is for, how it is computed, what good looks like, and where the reading was wrong before the counting was fixed. The figures come from one live deployment on a market research publisher's catalogue.
Every single metric for search has a way of being satisfied while search gets worse.
So the measures come as a set spanning three layers: what the engine does with queries you control, what it does with the queries real people type, and what that changes for the business.
Offline measures are the only ones readable on day one, and the only ones that say whether a change made search better or merely different.
Take a sample of items you know exist. For each, generate the ways a real person would ask for it badly: a letter deleted, two letters swapped, a word dropped, the words reordered, an abbreviation, a plural flipped, the subject without the catalogue's title convention, a natural-language phrasing. Run every variation and record whether the intended item appears in the top five and top ten.
Report the result per class, not only overall. The baseline on that catalogue is 450 sampled titles producing about 7,100 test queries per random seed, scored across 20 classes. Overall top-5 is 98.8%. The weakest class, abbreviations, began around 90% top-5, and that one row drove more engineering than the overall figure did.
The rule: a class dropping while the overall figure holds is a regression. A change that improved long queries by a few points quietly cost the drop-a-word class, and the overall number moved by less than the noise.
Building the test set, including the class list, is covered in a realistic dataset for comparing search solutions.
Generated variations test what you imagined. Historical failures test what happened. Every query from the logs that returned nothing, or that a customer complained about, goes into a second corpus with the item it should have found, versioned, and runs on every change alongside the generated set.
Good is 100% on this set and a steadily growing count of cases.
A search that never returns zero returns something for nonsense. Keep a set of queries that should return nothing: made-up words, phrases that combine two unrelated subjects. Score the number of results returned. Good is zero. This matters more once any meaning-based or fuzzy matching is in the path, because those components tend to answer everything.
Measure the engine on its own, per class, at p50 and p95. On the catalogue above most classes answer in 6 to 20 milliseconds at the median. Recording this offline is how a ranking improvement that costs time gets noticed before it meets the live budget: one relaxation change pushed title-shaped queries to 24 milliseconds, which was fine, and a later one pushed hyphenated terms past 300 milliseconds on the production box, which was not.
The share of real searches that returned nothing. The denominator matters: count intents, one query within a ten-minute window, not events. The first version of the console counted events and overstated searches by about a quarter.
What good looks like depends on where you start. The catalogue began at 53.4% zero results with the previous search. A random sample of 500 of those failures found that 470 had a relevant existing item, so the achievable floor was low. In the first days after the new engine went live, 5.1% of search intents returned nothing. The two figures have different denominators, raw searches before and intents after, and on either basis the fall is an order of magnitude.
Split the honest zeros out and keep them. A query that genuinely has no answer in the catalogue is demand.
Zero or non-zero is too coarse. Label each live search: zero, weak (results came back but none shares a subject word with the query), ok, or rescued (answered by a secondary channel such as spelling correction or meaning-based matching after the primary matching found nothing). The weak label finds the cases that embarrass: a three-word query that returned unrelated reports because the engine relaxed to its most common word. That one was found by reading the explorer, and fixed by making the rarest word in a query mandatory.
Good is a weak share that falls week on week and a rescued share you understand.
Every swallowed error in the search path is a counter on the health endpoint: fallbacks (the proxy gave up on the new engine and served the old one), degraded searches (a component failed but the search still answered), dropped telemetry, dropped clicks. Good is all zero. Non-zero counters are the difference between "search seems fine" and "search was fine for everyone except the people who typed hyphenated terms during the hour the timeout was too short."
Treat the timeout as a measure in its own right. The type-ahead budget on this deployment started at 300 milliseconds and moved to 600 after live data showed legitimate queries being cut off into fallback.
Type-ahead and a results page are different products with different budgets, so report them separately, at p50 and p95, per day, with the sample count next to every percentile. On this deployment the results-page rows carried no latency for the first week, because the path that logged them never recorded a timing; the tiles showed the type-ahead figures, which were excellent. A missing measure looks like a good one unless the sample size is printed.
A click is only meaningful when it can be joined to the search that produced it. Give every search an identifier when it is answered, hand it to the page, and send it back with each click together with the position clicked and the surface. Clicks that arrive with an identifier the engine never issued are kept but never counted, which also keeps automated traffic out of the click figures.
Then read three things: click rate per surface, position of the click, and searches that showed results nobody clicked. The last is the clearest "results did not convince" signal available.
The share of items that search surfaced at all in a period. In a sampled week before the change, the previous search had surfaced under 15% of the catalogue. Testing the corrected engine against the sample of failed queries alone surfaced more than a thousand reports that had not appeared in any result that week. A catalogue that cannot be reached through search is inventory being paid for and not sold.
Once zero results are mostly honest, the list of what people searched for and did not find is a product signal. Rank it by distinct intents. The console shows it for any window, and the publisher reads it as a commissioning list.
The measure everyone wants first and should take last, because it needs the other eleven to be interpretable. Join searches to the actions that matter in your business: an item opened, a sample requested, an order, a ticket resolved. Compare searches that converted with searches that did not, by quality label and by click position. Expect weeks of data before per-query figures mean anything on a site with modest search volume; aggregate figures are useful sooner. The conversion events live in commerce or service systems, not in the search engine, so this is an analytics join rather than a search feature; ITMTB's data and AI practice treats it that way.
| # | Measure | Layer | What good looks like | Read it |
|---|---|---|---|---|
| 1 | Known-item accuracy, per class | Offline | >98% top-5 overall; no class falls on a change | Every change |
| 2 | Real failed queries replayed | Offline | 100%, corpus growing | Every change |
| 3 | Nonsense guardrail | Offline | 0 leaks | Every change |
| 4 | Engine latency p50 / p95 per class | Offline | Well inside the live budget | Every change |
| 5 | Zero-result rate, by intent | Live | About 5%, remainder honest | Daily |
| 6 | Quality label (zero / weak / ok / rescued) | Live | Weak share falling | Weekly |
| 7 | Fallback, degraded, dropped counters | Live | All zero | Continuously |
| 8 | Latency per surface vs budget, with sample size | Live | p95 inside budget, no empty samples | Daily |
| 9 | Attributed clicks, per surface and position | Live | Unclicked-with-results share falling | Weekly, after a few weeks |
| 10 | Catalogue reach | Business | Rising share of items ever surfaced | Monthly |
| 11 | Demand in honest zeros | Business | A ranked list someone acts on | Monthly |
| 12 | Downstream conversion from search | Business | Converting share rising by label | Monthly, after a quarter |
When evaluating a tool rather than running one, measures 1 to 4 can be run on your own data before any contract, and the console that shows 5 to 9 should exist on day one.
ITMTB runs the offline layer against a prospective customer's catalogue and real failed queries before scoping a build; the offer is described on the enterprise search page. Contact us.
A person types "car" in the header, sees the dropdown, presses Enter and lands on a results page. That is one search and two events, sometimes from two IP addresses because the type-ahead travelled over IPv4 and the page request over IPv6. In one week's sample, most results-page rows had a type-ahead twin within ten minutes and few shared its IP; about a quarter of events were being counted twice. The fix was to define a search as a distinct query within a ten-minute window and to drop the IP from that key.
Bots hit search endpoints. They do not pick up the identifier the engine issues and send it back with a click. So search counts carry some noise and click joins carry none, which is one reason to read click rate as a ratio with care and the unclicked-with-results list with more confidence.
The same test set scored 98.8% one run and 98.6% the next with nothing changed in the engine. After an index rebuild, items with identical relevance scores came back in a different order, and a few cases sat exactly on the top-5 boundary. Treat differences below the noise as noise, record the index build with each score, and sort ties deterministically if the figure needs to be stable to a decimal.
Once a meaning-based channel answers queries the lexical engine could not, the zero-result rate falls. That is good only if the rescued answers are right. Label them, count them separately, and read them in the explorer weekly. Clicks on rescued results, once click attribution exists, are the evidence that the channel is worth its cost.
The results page had no latency figure for a week, and the percentile tiles showed the type-ahead numbers. Nothing was wrong except the reading. Every percentile now carries its sample size, and a surface with zero samples shows a dash.
| When | Do | Decide |
|---|---|---|
| Before launch | Build the known-item set from your items and your failed queries. Run measures 1 to 4. Fix anything under 95% in a class. Set latency budgets per surface. | Go or no-go on retrieval quality |
| First week | Read counters hourly on day one, then daily. Read every weak and rescued search by hand. Confirm intent dedup and sample sizes on the console. | Timeout and relaxation adjustments |
| Rest of the month | Add every real failure to the replay corpus. Start click attribution on every surface. Read zero-result and weak shares weekly. | Which failure classes need engine work and which a business rule |
| Second month onward | Catalogue reach, honest-zero demand, conversion join. Clicks become test cases and, bounded, a ranking signal. | Where search is leaving money |
The same measures serve a search used by AI agents as a retrieval tool, with one addition: track how often the agent's query returned nothing or something weak, because the agent reasons on whatever it gets back and the failure looks like a model problem. That dependency is covered in AI agents in business workflows.
There is no single one. The pair that works is a known-item test score measured offline, together with the live zero-result rate deduplicated by intent. Either alone can be gamed; together they test relevance and coverage.
Around 5% of distinct intents is achievable on a well-indexed catalogue. One catalogue went from 53.4% to 5.1% in the first days live. The remainder should be honest zeros, which are demand signals rather than failures.
With an offline evaluation: items you know exist, realistic variations of how people ask for them, scored top-5 and top-10 per variation class on every change, with any class dropping treated as a regression.
One person's search produces many events: a type-ahead request per keystroke burst, then a results-page request, sometimes from two IP addresses. Counting events overstated searches by roughly a quarter. Count distinct intents.
Yes, per surface, and only once a click can be attributed to the search that produced it. The most useful click signal is searches that returned results nobody clicked.
A few weeks on a site with modest search volume for aggregate measures. Per-query click rates stay noisy longer.
Tool metrics describe retrieval in isolation: offline accuracy, latency, corrections. Deployment metrics describe the tool inside your traffic: zero-result rate, fallbacks, counters, index freshness, clicks. A good tool can be a poor deployment.
The deployment behind these figures is ITMTB's enterprise search for The Business Research Company, described in the case study.
ITMTB runs the offline layer first: a known-item test set generated from your catalogue and your real failed queries, scored per class against your current search, before any build is scoped.