Between virtual machines, microservices, and AI pipelines, hybrid clouds can be incredibly complex and can bring an unwelcome partner: alert fatigue. SREs and IT OPs teams face a constant flood of disconnected alerts, forced to manually stitch together metrics from completely different monitoring tools. Sifting through raw logs and writing complex PromQL queries just to build static dashboards isn't sustainable.Something has to change. It's time to move away from fragmented tools and toward natural language and visualizations where the platform actually shows you the issues and helps you get the work done. If you're running modern enterprise workloads, then you're sitting on a goldmine of operational data — you just can't get to it easily. That's why we are expanding the cluster management capabilities of Red Hat OpenShift Lightspeed by combining it with new companion platform tools, such as the cluster observability operator.What is the cluster observability operator?In Red Hat OpenShift, an operator automates how Kubernetes-native applications are created, configured, and managed. The cluster observability operator introduces three key advancements that work together to help small teams scale with big clouds, turning day-to-day troubleshooting from a manual scavenger hunt into a simplified, intuitive workflow:Cluster incident detection: Adds an additional tab to the Observe > Alerting menu to help cluster operators reduce alert fatigue by automatically grouping and correlating related alerts into unified incidents.Signal correlation for Red Hat OpenShift: Speeds up troubleshooting by connecting your metrics, logs, alerts, and netflows across data stores into a single interactive graph, making data exploration more simple and direct.Red Hat build of Perses: A complete visualization overhaul that brings GitOps-native, custom dashboards-as-code directly into the OpenShift console, enabling the display of real-time metrics visualizations.In the rest of this post, we'll look at how these integrations work, and how they can help you move from constant firefighting to proactive, simplified cluster management.Why fragmented tools fail when seconds countHere's what we know from years of managing production incidents: When systems are down and every second counts, context switching kills response time. Jumping between separate monitoring dashboards, chat interfaces, and diagnostic tools doesn’t just slow you down –it fractures your understanding of the problem. Losing the thread and missing correlations leads to wasting precious minutes re-establishing context every time you switch tools.That’s why single-pane-of-glass operations are essential for effective incident management. OpenShift provides a unified view where natural language investigation, real-time metrics, and diagnostic insights live together. Red Hat build of Perses, signal correlation, incidents, and OpenShift Lightspeed interfaces are fully integrated into the OpenShift web console itself. No external tools. No silos. Just a cohesive troubleshooting experience where every component you need is right where you're already working: Your OpenShift cluster.From alert storm to root cause in minutesIt's Friday afternoon, and your monitoring dashboard lights up: Pod restarts are spiking in your production namespace, memory alerts are firing across different clusters, and your #ocp-incidents chat is scrolling faster than you can read. You know what that means: Hours of tedious troubleshooting.The old way: You hunt for the pod in the console, switch to your terminal for oc commands, jump to Prometheus to write PromQL queries, cross-reference metrics across separate dashboards, and grep through logs to manually piece together a timeline. Thirty minutes later, you might have a hypothesis.The modern streamlined way: You stay in the OpenShift console. With the cluster observability operator and OpenShift Lightspeed already running, you have everything you need to investigate in one place.Red Hat build of Perses dashboards, signal correlation, and incident detection capabilities are enabled in the cluster observability operator through UIPlugin custom resources. Red Hat build of Perses brings telemetry dashboards directly into your cluster and signal correlation provides a simplified, interactive view of your tied cluster resources. Finally, the cluster observability operator incident detection feature helps cluster operators to identify incidents displayed through severity-coded visual timelines. By organizing alerts around affected components and prioritizing them by impact level, you can rapidly trace problems to their source.
Curing alert fatigue: How embedded AI is redefining Red Hat OpenShift cluster troubleshooting
Redefine Red Hat OpenShift cluster troubleshooting with AI-powered tools. Simplify day-to-day workflow.
Red Hat automates incident correlation and log-metric alert unification in OpenShift Lightspeed, consolidating fragmented monitoring into single dashboard-as-code. SRE teams cut MTTR; hybrid cloud complexity scales without proportional cognitive overhead.






