We found out our webhook delivery was broken the same way most teams do: a customer told us. Not an alert, not a dashboard going red — an email asking why a payment status update from three hours earlier had never shown up in their system. By the time we'd pulled the logs, we counted just over six hours where a downstream endpoint had been failing silently, and every event meant for it was gone. Not queued, not retried. Gone.
The endpoint in question belonged to a mid-sized integration partner whose ops team had rotated a TLS cert and, in the process, briefly served a handshake our client library didn't like. Our webhook sender treated that as a delivery failure, logged it at a level nobody was watching, and moved on. There was no retry logic — we'd built a fire-and-forget publisher because in two years of running it, deliveries had basically always succeeded. That streak was the problem. It meant we'd never had to think about what "basically always" leaves out.
The fix looked simple on a whiteboard and took us about three weeks to get right in production, mostly because the interesting failure modes only show up under real load and real partner flakiness.
Exponential backoff, not fixed intervals






