UK hotel metasearch: 300 million queries a day at 80 percent lower infrastructure cost
A hotel availability search engine for a major UK bookings aggregator had reached the limit of a SQL Server cluster: adding a machine raised cost by 2 percent and throughput by 0.1 percent. We moved availability search into an in-memory data grid. Traffic grew 50 percent, response time fell 60 percent, infrastructure cost fell 80 percent.
- 300,000,000
- queries a day after the redesign, up from 200 million50% more traffic than the system it replaced
- -80%
- infrastructure cost, mostly persistent database usage removedProduction figure
- 95%
- of queries answered in under a second; response time down 60%Production figure
- Client
- A major UK hotel bookings aggregator
- Sector
- Travel: hotel metasearch and availability
- Workload
- About 200 million queries a day before the redesign, 99 percent read-only; 300 million after
- Stack
- Java, in-memory data grid for availability search, custom fast loader from the database, message queue pipeline for intraday updates, SQL Server retained as the system of record
- Offer today
- Performance Rescue and Managed Java Application Support
Scaling that cost 2 percent per machine and returned 0.1 percent
The search engine ran on database-oriented processing over a cluster of dozens of SQL Server machines. The architecture had stopped scaling horizontally: there were no options left to scale vertically, and adding a new SQL Server machine increased infrastructure cost by about 2 percent while raising throughput by about 0.1 percent.
Technical limits had become a business limit. Adding more clients and more hotel data was no longer cost-efficient, and the platform's growth stalled with it.
Availability search extracted into an in-memory data grid
We studied the detailed usage scenarios and statistics with the client and extracted availability search into a separate cache layer on an in-memory data grid. The grid holds the complete set of data required to answer a hotel availability query, so the request path never touches the database.
Two supporting components made that safe: a custom, fast loading mechanism to populate the grid from the database, and a message queue pipeline carrying incremental intraday updates so the in-memory view stays current. The database stayed as the system of record; it simply stopped being in the way of every query.
Measured after the first years in production
- -80%
- infrastructure cost, with persistent database usage sharply reducedProduction
- ~70%
- horizontal scaling efficiency: adding nodes now adds throughput almost linearlyProduction
- +50%
- traffic absorbed, to over 300 million queries a day, with response times down 60%Production
- Revenue
- from existing customers grew: they could query more often and generate higher volumesClient
The offers this engagement proves
Book a scoping call
Thirty minutes with an architect, not a salesperson. You leave with a written view of scope, price model and whether we are the right team for it.