Pod Pending: "Insufficient memory" — Requests, Node Pressure, or Limits Lie

The scheduler skips nodes that lack free memory for your pod's *request*. Sometimes the cluster is truly full; often it's oversized requests, unbalanced pools, or one leaky deployment eating the headroom.

What you'll see

Root causes

Requests larger than any node's allocatable

A 12Gi request on 8Gi nodes never schedules anywhere. Compare request vs kubectl describe node Allocatable: single big pods need big-node pools.

Allocatable is smaller than you think

The OS, kubelet, and eviction thresholds reserve memory — '8GB node' offers maybe 6.5GB allocatable. Other pods' requests (not usage!) consume it permanently.

One deployment's requests ballooned

kubectl describe node | head shows per-pod requests. A high-replica deployment with padded requests silently starves everything else of scheduling room.

Fix it

  1. Read the actual scheduling math
    kubectl describe node <n> | grep -A8 'Allocated resources' ; kubectl get pods -A -o jsonpath='{range .items[*]}{.metadata.name}{"\t"}{.spec.containers[*].resources.requests.memory}{"\n"}{end}' | sort -k2 -h | tail
  2. Right-size requests to observed usage
    kubectl top pods -A | sort -k3 -h | tail -10   # set requests near p95 usage, keep limits higher for bursts
  3. Scale down or split the request hog
    kubectl scale deployment <hog> --replicas=<n>   # or reduce replicas × request via HPA bounds
  4. Add nodes / bigger nodes for genuinely big pods
    # pods > (node allocatable / 3) won't pack efficiently; a dedicated larger pool beats request-shaving

Field note

Scheduling counts REQUESTS, not live usage — kubectl top showing free memory is irrelevant to a Pending 'Insufficient memory' pod. Consider a LimitRange + default requests so unconfigured deployments don't silently grab huge defaults.

Common questions

Node shows 4GB free but my 2GB pod won't schedule. Why?

Free RAM isn't schedulable RAM: the kubelet reserves some, and other pods' requests already claim the allocatable. Check describe node's Allocated resources — that's the ledger the scheduler uses.

Should requests equal limits?

For latency-critical services with stable usage, requests=limits (Guaranteed QoS) is reasonable. For bursty apps, modest requests + higher limits pack better; 100% requests everywhere wastes cluster money.

Ship it right the first time

Kustomize base with probes, PDBs, and zero-downtime rollouts already wired.

Kubernetes Production Blueprints — $27 →

One-time. Yours to modify. Instant download from the NinjaOps template store.