If you're a software engineer, architect, engineering manager, or platform engineer, I consider the Google SRE Book to be one of the handful of books that fundamentally changes how you think about running production systems. It's available for free online: Google Site Reliability Engineering Book.

Unlike many infrastructure books, it isn't about Kubernetes, AWS, or a particular technology. It's about the engineering principles behind operating systems at massive scale.

What makes it different?

Google's definition of SRE is: "What happens when you ask a software engineer to design an operations team."

Instead of treating operations as manual work, the philosophy is: