by Adam Gurary, Sheng Zhan, Ankit Vij, Vadim Antonov, Yu-Ju Huang and Dima Kotlyarov

Search is everywhere: product discovery on retail sites, voice queries on smart TVs, recommendations in every feed, and identity matching during account lookups. Each fires a request per page view, per keystroke, per user action. At consumer scale, that translates into thousands of queries per second hitting the search index, with peak traffic often several times higher.

Search also sits on the critical path of revenue. In retail, shoppers who use search convert at two to three times the rate of those who simply browse. On streaming platforms, recommendations drive most of what users watch.

Reaching production-level QPS used to mean building a bespoke retrieval stack per app. Load testing across index sizes, query types, and filters. Manually sizing capacity. Wiring up load balancers in front of the index. Every new search use case restarts the work. The pain is universal across vector databases, search engines, and DIY stacks.

Today we're announcing high-QPS scaling for Databricks AI Search which is generally available. Standard endpoints can now scale to thousands of QPS with a single, human-readable parameter. You tell us your target. We provision the infrastructure to meet it.