Originally published on kuryzhev.cloud
Your Grafana dashboard shows hundreds of failed SSH logins per minute from a single IP. The graph spikes at 3 a.m. like clockwork. And yet — nobody's phone buzzes. No PagerDuty ping, no Slack message, nothing. This is exactly the failure mode we hit last quarter while running Loki SSH brute-force detection across a fleet of bastion hosts, and it took longer than I'd like to admit to figure out why the pipeline was silently blind despite the data being right there in Loki.
Symptoms
The first sign wasn't an incident — it was a curious engineer poking around in Grafana Explore after a routine audit. They ran a quick query against the auth logs and saw a clear pattern: bursts of sshd[...]: Failed password for entries, clustered from a handful of rotating source IPs, mostly hitting during off-hours. Classic credential-stuffing behavior against SSH. Nothing subtle about it.
What was subtle: zero alerts had fired. No rule had triggered, no receiver had been paged, and the on-call rotation had no idea this had been happening for days. We went to the host directly — journalctl -u ssh showed hundreds of Invalid user and Failed password lines per minute during the attack windows. But the equivalent LogQL query in Loki returned either zero results or a partial subset, missing most of the actual volume.






