The scheduler skips nodes that lack free memory for your pod's *request*. Sometimes the cluster is truly full; often it's oversized requests, unbalanced pools, or one leaky deployment eating the headroom.
A 12Gi request on 8Gi nodes never schedules anywhere. Compare request vs kubectl describe node Allocatable: single big pods need big-node pools.
The OS, kubelet, and eviction thresholds reserve memory — '8GB node' offers maybe 6.5GB allocatable. Other pods' requests (not usage!) consume it permanently.
kubectl describe node | head shows per-pod requests. A high-replica deployment with padded requests silently starves everything else of scheduling room.
kubectl describe node <n> | grep -A8 'Allocated resources' ; kubectl get pods -A -o jsonpath='{range .items[*]}{.metadata.name}{"\t"}{.spec.containers[*].resources.requests.memory}{"\n"}{end}' | sort -k2 -h | tail
kubectl top pods -A | sort -k3 -h | tail -10 # set requests near p95 usage, keep limits higher for bursts
kubectl scale deployment <hog> --replicas=<n> # or reduce replicas × request via HPA bounds
# pods > (node allocatable / 3) won't pack efficiently; a dedicated larger pool beats request-shaving
Scheduling counts REQUESTS, not live usage — kubectl top showing free memory is irrelevant to a Pending 'Insufficient memory' pod. Consider a LimitRange + default requests so unconfigured deployments don't silently grab huge defaults.
Free RAM isn't schedulable RAM: the kubelet reserves some, and other pods' requests already claim the allocatable. Check describe node's Allocated resources — that's the ledger the scheduler uses.
For latency-critical services with stable usage, requests=limits (Guaranteed QoS) is reasonable. For bursty apps, modest requests + higher limits pack better; 100% requests everywhere wastes cluster money.
Kustomize base with probes, PDBs, and zero-downtime rollouts already wired.
Kubernetes Production Blueprints — $27 →One-time. Yours to modify. Instant download from the NinjaOps template store.