As asked
Your on-call rota is getting paged 40 times a week and people are burning out. How do you bring the load down without missing real incidents?
Sample answer outline
Treat alert hygiene as a programme, not a one-off. Categorise the last 4 weeks of pages: actionable vs auto-resolving, real impact vs symptom of a deeper issue, owned vs not owned. Delete alerts that auto-resolve and have no follow-up action. Merge duplicate alerts. Move symptom-based alerts to SLO burn-rate alerts. Push noisy alerts back to the team that owns the underlying system. Set a target alert budget per week and hold teams to it. Track it in retros.
Expect these follow-ups
- How do you stop a team from re-adding alerts you just removed?
- Which team owns an alert that fires from a shared system?
- What is the failure mode of burn-rate alerts?