If you're a software engineer, architect, engineering manager, or platform engineer, I consider the Google SRE Book to be one of the handful of books that fundamentally changes how you think about running production systems. It's available for free online: Google Site Reliability Engineering Book.
Unlike many infrastructure books, it isn't about Kubernetes, AWS, or a particular technology. It's about the engineering principles behind operating systems at massive scale.
What makes it different?
Google's definition of SRE is: "What happens when you ask a software engineer to design an operations team."
Instead of treating operations as manual work, the philosophy is:






