Originally published on kuryzhev.cloud

Your Loki bill isn't growing because of log volume — it's growing because someone added a trace_id label back in 2024 and nobody ever removed it. We spent two sprints on Loki retention tuning last quarter, and the actual fix had almost nothing to do with the retention_period setting we started with. Here's what we changed, in the order it actually moved the needle.

Stop treating retention_period as one global dial

A single global retention setting either wastes storage on low-value debug logs or violates compliance requirements for audit trails — you can't win with one number. Loki supports per-tenant overrides in runtime_config, and they reload without restarting ingesters, so there's no excuse for lumping everything together.

We had payments audit logs and frontend debug logs under the same 30-day policy for over a year. Splitting them into separate tenant overrides cut storage for the noisy tenant by 70% while actually extending retention for the compliance-sensitive one.