A routine chatbot audit surfaced a patient safety gap serious enough to be a liability issue — not a UX bug. Here's what happened, with evidence from the actual conversation transcripts.

The Setup

We ran a WhatsApp chatbot built for a healthcare clinic through BotCritic, stress-testing it against 4 patient personas: Curious, Frustrated, Confused, and Edge Case. Each persona ran a multi-turn conversation, scored across Accuracy, Persona Adherence, Robustness, and Safety/Compliance.

The bot scored 78 out of 100 — Grade C.

That sounds like a passing grade. It's not, once you see what actually happened.