Kubernetes node stuck NotReady: diagnosis order that works
⏱️ 2 min read
What NotReady means
The node object's lease stopped being renewed, so the scheduler stops placing pods and existing pods get evicted after the grace window. The cause is almost always kubelet itself: crashed, paused, or unable to reach the API server.
Diagnosis order
- Node conditions and events.
kubectl describe node <name>:KubeletNotReadymessages tell you the sub-reason (PLEG is not healthy, network plugin not ready, disk pressure). - Kubelet logs on the node.
journalctl -u kubelet -n 100 --no-pager. The single most informative step. - Container runtime.
systemctl status containerd(or docker). A dead runtime takes kubelet down with it. - Disk.
df -h /var/lib/kubelet /var/lib/containerd. Image garbage collection stops, then the whole node wedges. The fix is cleanup, not more pods. - Network plugin. CNI daemonset crashlooping shows
networkPluginNotReadyforever. - Certificates. Client cert expiry gives 401s in kubelet logs; kubelet cannot renew certs if the clock is skewed (check
timedatectl).
Quick wins
Restart kubelet only after reading logs; a blind restart hides the cause and it comes back. If PLEG is the message, check containerd first, PLEG is downstream of runtime health.
For a lab where you can afford to break nodes on purpose, a managed cluster costs cents per node-hour: DOKS or a Vultr VM with kubeadm. (Partner links.)