The Business Research Company, a global market research publisher with 25,700 published reports

Enterprise Search Case Study: How ITMTB Took The Business Research Company's Zero-Result Searches From 53% to 5%

Outcome

Zero-result searches down from 53.4% to 5.1% of search intents, 98.8% known-item accuracy on the live catalogue, type-ahead p50 25 ms

Enterprise Search Case Study: How ITMTB Took The Business Research Company's Zero-Result Searches From 53% to 5%
By Aakash Ahuja2026-10-03

The short version

The Business Research Company publishes more than 25,000 market research reports and sells them through its website. On that site, enterprise search is the path to revenue: a buyer knows the market they need, types it, and either finds the report or concludes it does not exist. This case study covers what was wrong, what ITMTB built, how it was measured and what changed.

In a sampled week, 53.4% of searches returned nothing. A random sample of 500 of those failed searches, checked against the catalogue by hand, found that 470 had a relevant report that already existed. The content was there. Search was the barrier.

ITMTB designed, built and launched a replacement search engine tuned to the catalogue, measured it on the catalogue's own titles before it shipped, and put it into production within three weeks of design sign-off, with the previous search as a fallback at every step. In the first days on the live site, 5.1% of search intents returned nothing.

The engineering detail is in three companion articles: the retrieval failures found in the logs, how search success is measured and the test set that gates every change.

Why could buyers not find reports that existed?

ITMTB already ran the customer's website operations; that engagement is described in an earlier case study, and it is how search came into view. The site had become fast and reliable. Search had not become good.

A week of search logs told the story:

  • 53.4% of searches returned no result.
  • Hundreds of sequences of three or more consecutive failures were people reformulating and failing again.
  • Under 15% of the catalogue appeared in any result, for anyone, all week.
  • Zero results rose with query length, from 17.7% for one-word queries to about 69% for three or more words. The more precisely a buyer described what they wanted, the more likely they were to be told it did not exist.

The causes were small and specific, and none of them was "the catalogue lacks it". Words such as "latest" and "report" were treated as mandatory. Every added word was another way to fail. A general dictionary corrected valid specialist terms into other words. Two-character terms such as AI, 5G and EV were absent from the index because of a database default.

The business saw a catalogue of 25,000 reports. The buyer saw a search box that said no.

Why was "add AI search" the wrong frame?

The current reflex is a conversational search box. The reason to resist it is easy to test: if retrieval returns the wrong subset, a language model produces a confident answer about the wrong report, and the failure is blamed on the model. Retrieval has to work first.

So the engine's job was kept narrow and testable: return the most relevant existing reports for an imperfect query, fast. Generation was left for jobs that need it. A meaning-based channel was built, but gated so that it runs only when lexical matching finds nothing and only above a confidence floor calibrated so nonsense still returns nothing.

That is the same discipline ITMTB applies to agentic AI: give each layer a job it can be measured on.

What did ITMTB build?

A domain-tuned search engine as a separate service. The engine reads the report database into its own index, in its own schema, and runs as its own process with its own database user and connection pool, so nothing search-side can touch orders or enquiries. The existing website backend calls it over a local proxy with a timeout and a fallback to the old search. The search box and results page the visitor sees did not change.

Retrieval tuned to the catalogue, not to a dictionary. Spelling correction draws its vocabulary from the catalogue's own titles, so "virtopsy" is not corrected into something else. Convention words that carry no subject are demoted using the catalogue's own frequency, not a generic stop-word list. Longer queries relax progressively while the rarest word stays mandatory, so adding words helps instead of hurting. Short terms such as AI and EV are indexed. Hyphenated terms such as CAR-T are kept whole.

An evaluation harness on the customer's catalogue. Before anything shipped, 450 report titles were sampled and about 7,100 realistic queries generated from them in 20 variation classes: five kinds of typo, dropped and extra words, three kinds of reordering, abbreviations, plurals, hyphens, natural phrasing and the catalogue's title convention. The engine must return the intended report in the top five; it does for 98.8%, and every class is tracked on every change. Six further classes were added from live complaints after launch.

A search console. Every search is recorded with its results, a quality label (zero, weak, ok or rescued), latency, engine version and the clicks it produced. The console shows zero-result queries ranked by demand, latency percentiles per surface with sample sizes, error counters, and a query explorer for "a customer says they searched for X". Counting is by intent, one query within ten minutes, so a person typing is not five searches.

Business rules. The customer's team can pin, exclude or redirect specific reports for specific queries from the console, with a preview, without a deployment.

The same search for AI assistants. The engine is exposed through the Model Context Protocol, so an assistant asking for reports on a market gets what the website's search box would show. The commercial case is in MCP for business.

Operations built in. A nightly index rebuild with an incremental refresh every fifteen minutes, a drift check that compares the index with the published catalogue on every health read, counters for every swallowed error, alerts when the index falls silent, and a nightly incremental embedding refresh for the semantic channel. Deployment is through a tagged release pipeline with manual approval; rollback to the old search is one configuration value.

What does the engine do now, improvement by improvement?

Each of these came from a measured failure, in the logs or from the customer's own testing, and each is a committed test case the engine must keep passing.

ImprovementWhat the buyer seesWhere it came from
Convention words demoted by catalogue frequency"latest report on carbon capture" finds the carbon capture report46% of the sampled failures
Progressive relaxation with the rarest word kept mandatoryAdding words narrows instead of killing the search; "robot joint position sensor" can no longer return drillshipsLong-query failures; one live complaint
Spelling correction from the catalogue's own vocabulary, with a vowel-swap channel"catalist" becomes catalyst; "virtopsy" is left aloneDictionary damage in the logs; a customer-reported term
Short and compound terms kept wholeAI, 5G, EV, Wi-Fi and CAR-T work; CAR-T never matches "car"Index default of three characters; customer testing
Exact-title and exact-subject promotionA query that is a report's subject puts that report first, every timeCustomer rule: an exact match is always on top
Country and year qualifiers"France pest control 2024" finds the pest control report; the country and year narrow rather than excludeCustomer testing; a synthetic set of such queries failed almost entirely before the fix
One engine for the header, the home grid and the results pagePressing Enter shows the same list the dropdown promised, with the true total and pagingThe old results page used a different, slower search
Typo rescue on the results pageA misspelt Enter search shows "Showing results for …" instead of nothingResults-page zero results
Business rules with previewThe customer pins, excludes or redirects reports for specific queries, instantlyCommercial priorities no ranking model should guess

How is semantic search used, and when is it not?

Semantic matching, where queries and reports are compared by meaning rather than by words, was built into the engine but not placed in front of the lexical path. It does two bounded jobs:

  • Rescue. When lexical matching, correction and rules all settle on zero, the query is embedded and compared with every report's embedding. A result is shown only above a confidence floor calibrated so that nonsense still returns nothing: "fancy expensive hotels" is rescued to the hotel reports, "blockchain basketweaving" stays empty.
  • Related reports. When a short query has one confident answer, the nearest reports to that answer are appended and labelled, so "CAR-T" shows the CAR-T therapy report first and then leukaemia, targeted therapy and immunotherapy reports, instead of seven unrelated titles.

Both run on a small open model inside the service, with no external API and no per-query cost; the embeddings are refreshed incrementally each night in seconds. The channel is a per-customer switch. On the test environment both jobs are on and verified; in production they switch on after the first embedding build, scheduled off-peak. Until then production answers lexically, which is what every accuracy figure in this case study measures.

How is success measured?

The engine is judged on three layers, described in full in how to measure the success of an enterprise search tool:

  1. Offline, on every change: known-item accuracy per variation class (98.8% top-5 across 20 classes, plus 6 classes from live complaints), the customer's real failed queries replayed, a nonsense set that must return nothing, and engine latency per class.
  2. Live, on the console: zero-result rate counted by intent, a quality label on every search, fallback and degraded counters, latency per surface with sample sizes, and clicks attributed to the search that produced them, from the dropdown and from the results page.
  3. Business, monthly: catalogue reach, demand visible in honest zero results, and the join from search to enquiries and orders.

The rule the customer signed off on keeps the figures honest: a variation class dropping while the overall score holds is a regression, and no change ships without its test committed first.

How did it reach production, and how fast?

StageWhat happened
AnalysisA week of search logs: 53.4% zero results; 470 of 500 sampled failures had an existing answer
DesignAgreed with the customer: separate service and schema, legacy search frozen as fallback, evaluation gate on every change
Same weekEngine ported and measured on the catalogue: 98.6% / 97.7% top-5 on two seeds
One week after designLive on the customer's test environment; the customer tested with their own queries and the ones their buyers had typed, and every finding became a committed test before it was fixed
Following weekCustomer feedback fixed; console filters and export added
Three weeks after designProduction: database, service pipeline, backend port and front end cut over; the new search serving all traffic
First days liveThree further releases from live findings: ranking fixes from the customer's own queries, results-page click tracking, and the related-reports and semantic channel built and verified on the test environment, switching on in production after the first embedding build

The cutover needed no downtime. Three properties were built in rather than assumed: the old search answers if the new service fails or times out; the service cannot reach the commerce tables; and every error the engine swallows is a counter that someone reads.

What changed for the customer?

MeasurePrevious search (sampled week)ITMTB engine (first days live)
Searches returning nothing53.4% of searches5.1% of search intents
Searches rescued by typo correctionnoneabout one in ten
Known-item accuracy, top 5 (offline, 20 classes)not measurable98.8%
Type-ahead latencynot measuredp50 25 ms, p95 247 ms
Fallbacks, degraded searches, dropped eventsnot visible0, 0, 0 (all counted)
Index coverageunknownevery published report, checked on every health read
Who can see what people search fornobodythe customer's team, live, with clicks

The two zero-result figures have different denominators: the old logs counted raw searches, the new console counts intents, which is a smaller number. On either basis the fall is an order of magnitude. Catalogue reach and search-to-order conversion are measured monthly and are not yet published.

One change the table does not show: the customer's team now reads the top zero-result queries as a commissioning list, what buyers asked for that the catalogue does not have.

Key takeaways

  • Most zero-result searches were retrieval failures: 470 of 500 sampled failures had an existing answer.
  • Fixing retrieval before adding any conversational layer took zero results from 53.4% to 5.1% of search intents.
  • Measurement came before the build: a per-class known-item test on the customer's own catalogue, 98.8% top-5, gating every change since.
  • A separate service with a fallback at every step went from design to production in three weeks with no downtime.
  • The same engine now serves the website, the console, business rules and AI assistants through MCP.

How do you recognise this kind of problem?

You probably have it if any of these are true:

  1. You do not know what share of your searches return nothing.
  2. You have never replayed a sample of failed searches to see how many had an answer.
  3. Searching for something you know exists, misspelt or with a word dropped, loses it.
  4. Longer, more specific queries do worse than short ones.
  5. The plan for fixing search is a conversational interface on top of it.

The fix is a retrieval project with a measure, not a platform replacement. What ITMTB deploys, and how an engagement runs, is on the enterprise search page; the data and AI practice scopes it against the customer's own catalogue first.

Want a similar outcome?

Tell us about your problem. We'll tell you if we've seen it before and how we'd approach it.

Start a conversation →

More Success Stories

Improving Enterprise Search for People and AI Agents

Improving Enterprise Search for People and AI Agents

Read More
How to Measure the Success of an Enterprise Search Tool: 12 Measures From a Live Deployment

How to Measure the Success of an Enterprise Search Tool: 12 Measures From a Live Deployment

Read More
A Realistic Dataset for Comparing Enterprise Search Solutions

A Realistic Dataset for Comparing Enterprise Search Solutions

Read More
Why Enterprise Search Is Important: The Cost of Content Nobody Can Find

Why Enterprise Search Is Important: The Cost of Content Nobody Can Find

Read More
Optimizing and Securing Website Operations for a Global Market Intelligence Business

Optimizing and Securing Website Operations for a Global Market Intelligence Business

Read More

Ready to Transform Your Business?

Join industry leaders already scaling with our custom software solutions. Let’s build the tools your business needs to grow faster and stay ahead.