The development of two medical AI assistants highlights an unnerving challenge: as the technology races ahead, what is the best way to evaluate what works?