The Business Research Company, a global market research publisher with 25,700 published reports
Outcome
Zero-result searches down from 53.4% to 5.1% of search intents, 98.8% known-item accuracy on the live catalogue, type-ahead p50 25 ms

The Business Research Company publishes more than 25,000 market research reports and sells them through its website. On that site, enterprise search is the path to revenue: a buyer knows the market they need, types it, and either finds the report or concludes it does not exist. This case study covers what was wrong, what ITMTB built, how it was measured and what changed.
In a sampled week, 53.4% of searches returned nothing. A random sample of 500 of those failed searches, checked against the catalogue by hand, found that 470 had a relevant report that already existed. The content was there. Search was the barrier.
ITMTB designed, built and launched a replacement search engine tuned to the catalogue, measured it on the catalogue's own titles before it shipped, and put it into production within three weeks of design sign-off, with the previous search as a fallback at every step. In the first days on the live site, 5.1% of search intents returned nothing.
The engineering detail is in three companion articles: the retrieval failures found in the logs, how search success is measured and the test set that gates every change.
ITMTB already ran the customer's website operations; that engagement is described in an earlier case study, and it is how search came into view. The site had become fast and reliable. Search had not become good.
A week of search logs told the story:
The causes were small and specific, and none of them was "the catalogue lacks it". Words such as "latest" and "report" were treated as mandatory. Every added word was another way to fail. A general dictionary corrected valid specialist terms into other words. Two-character terms such as AI, 5G and EV were absent from the index because of a database default.
The business saw a catalogue of 25,000 reports. The buyer saw a search box that said no.
The current reflex is a conversational search box. The reason to resist it is easy to test: if retrieval returns the wrong subset, a language model produces a confident answer about the wrong report, and the failure is blamed on the model. Retrieval has to work first.
So the engine's job was kept narrow and testable: return the most relevant existing reports for an imperfect query, fast. Generation was left for jobs that need it. A meaning-based channel was built, but gated so that it runs only when lexical matching finds nothing and only above a confidence floor calibrated so nonsense still returns nothing.
That is the same discipline ITMTB applies to agentic AI: give each layer a job it can be measured on.
A domain-tuned search engine as a separate service. The engine reads the report database into its own index, in its own schema, and runs as its own process with its own database user and connection pool, so nothing search-side can touch orders or enquiries. The existing website backend calls it over a local proxy with a timeout and a fallback to the old search. The search box and results page the visitor sees did not change.
Retrieval tuned to the catalogue, not to a dictionary. Spelling correction draws its vocabulary from the catalogue's own titles, so "virtopsy" is not corrected into something else. Convention words that carry no subject are demoted using the catalogue's own frequency, not a generic stop-word list. Longer queries relax progressively while the rarest word stays mandatory, so adding words helps instead of hurting. Short terms such as AI and EV are indexed. Hyphenated terms such as CAR-T are kept whole.
An evaluation harness on the customer's catalogue. Before anything shipped, 450 report titles were sampled and about 7,100 realistic queries generated from them in 20 variation classes: five kinds of typo, dropped and extra words, three kinds of reordering, abbreviations, plurals, hyphens, natural phrasing and the catalogue's title convention. The engine must return the intended report in the top five; it does for 98.8%, and every class is tracked on every change. Six further classes were added from live complaints after launch.
A search console. Every search is recorded with its results, a quality label (zero, weak, ok or rescued), latency, engine version and the clicks it produced. The console shows zero-result queries ranked by demand, latency percentiles per surface with sample sizes, error counters, and a query explorer for "a customer says they searched for X". Counting is by intent, one query within ten minutes, so a person typing is not five searches.
Business rules. The customer's team can pin, exclude or redirect specific reports for specific queries from the console, with a preview, without a deployment.
The same search for AI assistants. The engine is exposed through the Model Context Protocol, so an assistant asking for reports on a market gets what the website's search box would show. The commercial case is in MCP for business.
Operations built in. A nightly index rebuild with an incremental refresh every fifteen minutes, a drift check that compares the index with the published catalogue on every health read, counters for every swallowed error, alerts when the index falls silent, and a nightly incremental embedding refresh for the semantic channel. Deployment is through a tagged release pipeline with manual approval; rollback to the old search is one configuration value.
Each of these came from a measured failure, in the logs or from the customer's own testing, and each is a committed test case the engine must keep passing.
| Improvement | What the buyer sees | Where it came from |
|---|---|---|
| Convention words demoted by catalogue frequency | "latest report on carbon capture" finds the carbon capture report | 46% of the sampled failures |
| Progressive relaxation with the rarest word kept mandatory | Adding words narrows instead of killing the search; "robot joint position sensor" can no longer return drillships | Long-query failures; one live complaint |
| Spelling correction from the catalogue's own vocabulary, with a vowel-swap channel | "catalist" becomes catalyst; "virtopsy" is left alone | Dictionary damage in the logs; a customer-reported term |
| Short and compound terms kept whole | AI, 5G, EV, Wi-Fi and CAR-T work; CAR-T never matches "car" | Index default of three characters; customer testing |
| Exact-title and exact-subject promotion | A query that is a report's subject puts that report first, every time | Customer rule: an exact match is always on top |
| Country and year qualifiers | "France pest control 2024" finds the pest control report; the country and year narrow rather than exclude | Customer testing; a synthetic set of such queries failed almost entirely before the fix |
| One engine for the header, the home grid and the results page | Pressing Enter shows the same list the dropdown promised, with the true total and paging | The old results page used a different, slower search |
| Typo rescue on the results page | A misspelt Enter search shows "Showing results for …" instead of nothing | Results-page zero results |
| Business rules with preview | The customer pins, excludes or redirects reports for specific queries, instantly | Commercial priorities no ranking model should guess |
Semantic matching, where queries and reports are compared by meaning rather than by words, was built into the engine but not placed in front of the lexical path. It does two bounded jobs:
Both run on a small open model inside the service, with no external API and no per-query cost; the embeddings are refreshed incrementally each night in seconds. The channel is a per-customer switch. On the test environment both jobs are on and verified; in production they switch on after the first embedding build, scheduled off-peak. Until then production answers lexically, which is what every accuracy figure in this case study measures.
The engine is judged on three layers, described in full in how to measure the success of an enterprise search tool:
The rule the customer signed off on keeps the figures honest: a variation class dropping while the overall score holds is a regression, and no change ships without its test committed first.
| Stage | What happened |
|---|---|
| Analysis | A week of search logs: 53.4% zero results; 470 of 500 sampled failures had an existing answer |
| Design | Agreed with the customer: separate service and schema, legacy search frozen as fallback, evaluation gate on every change |
| Same week | Engine ported and measured on the catalogue: 98.6% / 97.7% top-5 on two seeds |
| One week after design | Live on the customer's test environment; the customer tested with their own queries and the ones their buyers had typed, and every finding became a committed test before it was fixed |
| Following week | Customer feedback fixed; console filters and export added |
| Three weeks after design | Production: database, service pipeline, backend port and front end cut over; the new search serving all traffic |
| First days live | Three further releases from live findings: ranking fixes from the customer's own queries, results-page click tracking, and the related-reports and semantic channel built and verified on the test environment, switching on in production after the first embedding build |
The cutover needed no downtime. Three properties were built in rather than assumed: the old search answers if the new service fails or times out; the service cannot reach the commerce tables; and every error the engine swallows is a counter that someone reads.
| Measure | Previous search (sampled week) | ITMTB engine (first days live) |
|---|---|---|
| Searches returning nothing | 53.4% of searches | 5.1% of search intents |
| Searches rescued by typo correction | none | about one in ten |
| Known-item accuracy, top 5 (offline, 20 classes) | not measurable | 98.8% |
| Type-ahead latency | not measured | p50 25 ms, p95 247 ms |
| Fallbacks, degraded searches, dropped events | not visible | 0, 0, 0 (all counted) |
| Index coverage | unknown | every published report, checked on every health read |
| Who can see what people search for | nobody | the customer's team, live, with clicks |
The two zero-result figures have different denominators: the old logs counted raw searches, the new console counts intents, which is a smaller number. On either basis the fall is an order of magnitude. Catalogue reach and search-to-order conversion are measured monthly and are not yet published.
One change the table does not show: the customer's team now reads the top zero-result queries as a commissioning list, what buyers asked for that the catalogue does not have.
You probably have it if any of these are true:
The fix is a retrieval project with a measure, not a platform replacement. What ITMTB deploys, and how an engagement runs, is on the enterprise search page; the data and AI practice scopes it against the customer's own catalogue first.
Tell us about your problem. We'll tell you if we've seen it before and how we'd approach it.
Start a conversation →Join industry leaders already scaling with our custom software solutions. Let’s build the tools your business needs to grow faster and stay ahead.