CrashLoopBackOff: the systematic fix path
⏱️ 2 min read
Read the last crash first
kubectl logs <pod> --previous kubectl describe pod <pod>
--previous is the important one: it shows the logs of the container that already died. kubectl describe shows the exit code and reason, plus whether a probe (liveness/readiness) killed it.
The five causes, in order of frequency
- App crashes on startup. Missing env var, bad secret, unreachable DB. The previous logs say so directly. Fix the app or config; nothing in Kubernetes needs to change.
- Liveness probe too aggressive. If the app needs 30s to boot but
initialDelaySecondsis 5 and the probe interval is 5, kubelet kills it mid-boot, forever. CheckLast State: Terminatedwith exit code 137 andReason: OOMKilledvs probe kills in events. - OOMKilled (exit 137). Raise
resources.limits.memoryabove peak usage, or fix the leak. - Command/args override mistakes. A wrong
command:in the manifest replaces the image's entrypoint and the container dies instantly with exec format errors. - Dependency ordering. The app dies because the DB is not reachable yet. Use init containers or readiness gates, not sleep loops.
Make it tell you earlier
Crash loops burn restart tokens for nothing. Add a startupProbe so slow booters get a grace window, and keep liveness for post-boot hang detection only. If you are tuning this repeatedly on a small cluster, running a disposable cluster for a few cents an hour beats breaking prod on a Friday: DOKS on DigitalOcean is the usual lab. (Partner link.)