Perplexity details its embedding serving stack: Ivy, Tulip and ROSE, CUDA graph capture, LazyTensor overlap and vLLM benchmarks.

Perplexity AI publishes research on custom serving infrastructure including ROSE engine and pplx-embed models to cut costs across 400 million

Perplexity details its embedding serving stack: Ivy, Tulip and ROSE, CUDA graph capture, LazyTensor overlap and vLLM benchmarks.