Kubernetes Fixes
21 troubleshooting guides — each with the error, why it happens, and copy-paste commands to fix it.
- Kubernetes Pod Stuck in CrashLoopBackOff — CrashLoopBackOff is not the error — it is Kubernetes giving up on restarting a container that keeps dying.
- Kubernetes Pod Stuck in ImagePullBackOff — The kubelet cannot pull your container image.
- kubectl: 'Connection Refused' on Port 6443 — kubectl cannot reach the API server.
- Kubernetes Pod Stuck in Pending (Nothing Is Wrong With It) — Pending means the scheduler cannot find a home for the pod.
- Kubernetes Ingress Returns 404 for Everything — The controller is up and answering — that is why you get a tidy 404.
- Kubernetes Node Stuck in NotReady — The API server remembers the node, but kubelet stopped reporting.
- Kubernetes Service Gives 'Connection Refused' to a Live Pod — Your app pod is Running and the Service exists — but connections refuse.
- Helm Release Stuck in 'pending-upgrade' — A failed upgrade left the release in limbo and Helm refuses to continue.
- Kubernetes: Weird TLS/Cert Errors Everywhere (Check the Clock) — Expired-looking x509 errors, 'token invalid', or nodes failing to join with intact certs is the classic symptom of clock skew.
- kubectl Returns 401 Unauthorized — The API server rejected your credential — expired token, stale kubeconfig, or a certificate issue.
- Kubernetes PVC Stuck in Pending — A PersistentVolumeClaim that never binds means no StorageClass default, a provisioning gap, or a topology conflict.
- Kubernetes: "Volume Node Affinity Conflict" (PVC Zone Pinning) — A PersistentVolumeClaim exists in one availability zone; the scheduler can only put pods on nodes in that zone.
- Pod Pending: "0/1 nodes are available: 1 node(s) had untolerated taint" — The scheduler found a node but refuses to use it: the node carries a taint your pod doesn't tolerate.
- Pod Pending: "Insufficient memory" — Requests, Node Pressure, or Limits Lie — The scheduler skips nodes that lack free memory for your pod's *request*.
- kubectl "Context Deadline Exceeded" — API Server Too Slow or Unreachable — kubectl gave up waiting on the API server.
- Kubernetes Pods Can't Resolve Service Names (CoreDNS/ndots Debug) — nslookup works for external names but service-name lookups fail — or everything DNS is flaky.
- Kubernetes ConfigMap Changed but Pods Don't See It — Volume-mounted ConfigMaps update (slowly, and only for some subPath cases); envFrom/Env vars NEVER update.
- Kubernetes Ingress 502 / "Connection Refused by Upstream" — Service Chain Debug — Ingress → Service → Endpoints → Pod: a 502 means the chain broke at one link.
- kubectl: "No Configuration Has Been Provided" — Kubeconfig Absent or Broken — kubectl can't find a kubeconfig with a current context: file missing, KUBECONFIG pointing nowhere, or context deleted.
- Kubernetes Evicting Pods: Node Under Memory/Disk Pressure — The kubelet evicts pods when a node runs low on memory or disk.
- kubectl: "x509: Certificate Signed by Unknown Authority" — kubectl doesn't trust the API server's certificate authority.