As asked
Tell me about a time you shipped a feature or fix that caused a production incident. What broke, how did you find out, what did you do to restore the service, and what did you change afterward so it did not happen again?
Sample answer outline
A strong answer walks through the specific incident with concrete details: what changed, what failed, the timeline from detection to resolution, and who was affected. The candidate should show they took ownership, communicated proactively, and focused on restoring service before fixing blame. The retrospective section should include a systemic change like adding a test, changing a deploy process, or adding a monitor, not just 'I was more careful.'
Expect these follow-ups
- How long did it take to detect the issue? What would have made detection faster?