Entrusting mission-critical workloads in the cloud puts a lot of pressure on IT operations to build resilience into data infrastructure they don’t even own — to make sure that when an incident happens elsewhere, the lights stay on at home. Fortunately, savvy CIOs are learning they can architect their systems for resilience and high availability to keep mission-critical cloud workloads running when disaster strikes.
When designing for resilient, highly available cloud workloads it is critical to understand the details of your cloud provider’s service level agreement (SLA) and recognize that providers operate on a shared responsibility model. That means there are many things the provider is responsible for, but the workloads you’ve got running on their infrastructure are not among them. With that in mind, we’ll need to focus on three core goals essential to building and maintaining highly available, resilient cloud workloads:
Identify and eliminate single points of failure
Close or mitigate conditions leading to recovery gaps, and
Overcome operational complexity.







