As asked
Describe a significant Kafka outage or incident you were responsible for resolving. Walk me through what failed, how you diagnosed it in real time, what you did to restore service, and what you changed afterward to prevent a recurrence.
Sample answer outline
Strong answers use a clear timeline, quantify the impact (messages lost, lag accumulated, SLA breach), name the specific Kafka metrics or broker logs that pointed to the root cause, describe the recovery steps taken under pressure (e.g., preferred leader election, increasing replica timeout, rolling restart), and articulate a concrete post-incident change (runbook, monitoring, config fix). Vague answers that say 'we fixed it' without specifics are weak.
Expect these follow-ups
- What monitoring gap allowed this incident to go undetected for as long as it did?
- If you had to do it again, what would you do differently in the first 15 minutes?