As asked
A Go service slowly grows from 200 goroutines to 50,000 during normal traffic. Walk me through how you would diagnose and fix the leak.
Sample answer outline
Start with evidence: runtime metrics, pprof goroutine dumps and stack traces grouped by blocked call site. Common causes are sends to channels nobody receives from, receives on channels nobody closes, missing context cancellation and background tickers that are never stopped. The fix should remove the lifecycle ambiguity, usually by passing context.Context through the call chain, closing owned channels exactly once, and stopping tickers or timers. The candidate should also discuss regression protection with load tests or a goroutine-count assertion around the failing path. Interviewers look for ownership of goroutine lifetimes, not just knowing pprof commands.
Expect these follow-ups
- How do you decide which side owns closing a channel?
- What does a goroutine blocked on select usually tell you?
- How would you expose goroutine growth as an alert?