How do you troubleshoot a Pod stuck in CrashLoopBackOff?
Assesses fundamental understanding of Kubernetes conventions, runtime behavior, and memory/performance considerations.
Hiring managers look for precision, avoidance of ambiguous jargon, and ability to explain trade-offs under real production conditions.
CrashLoopBackOff means the container starts, exits, and Kubernetes restarts it with exponential backoff.
- Get the state and previous logs.
kubectl get pod web-1 -o wide
kubectl describe pod web-1
kubectl logs web-1 --previous
- Read the exit code: 1 is an application error, 137 is OOMKilled, and 127 is command not found.
- Common causes:
- Application config or missing environment variables and Secrets.
- Failing readiness or liveness probes causing restarts.
- Out-of-memory: raise limits or fix a leak.
- Permission errors on mounted volumes.
- A wrong entrypoint or command.
- Debug interactively by overriding the command, then exec in and reproduce.
- Check events with kubectl get events --sort-by=.lastTimestamp and inspect the failing dependency.
Fix the root cause rather than only increasing the restart backoff.
Candidate Response Strategy & Interview Tips
- Start with a concise one-sentence summary: Deliver a direct, confident answer first before expanding into nuances.
- Demonstrate real-world trade-offs: Discuss where this approach excels and when you would avoid it in production systems.
- Discuss complexity & edge cases: Proactively explain time/space complexity or boundary conditions (null values, scale limits).
- Prepare for interviewer follow-ups: Technical hiring panels frequently probe deeper into concurrency, backward compatibility, or alternative libraries.