Perplexity AI is releasing WANDR, a benchmark designed to evaluate how well AI agents handle complex, multi-layered research tasks. The open-source release gives developers and researchers a standardized way to measure whether their AI systems can actually do serious investigative work.
WANDR, which stands for Wide and Nuanced Deep Research, is built specifically for what Perplexity calls “wide research” tasks that require broad searching and deep investigation, aligned with Perplexity Computer’s capabilities.
What WANDR actually measures
The benchmark is designed to mirror the demanding workloads typically found in professional research environments, going well beyond simple question-answering or single-document summarization.
Perplexity positions WANDR as an evolution of earlier evaluation frameworks like WideSearch, reflecting a shift towards more realistic professional workloads.






