Perplexity AI open sources WANDR, a benchmark for evaluating AI research agents on complex wide research tasks, with its SaC framework scoring 2.5x above