Kubernetes requests vs limits: the practical guide
⏱️ 2 min read
Why both numbers matter
requests is what the scheduler believes and reserves; limits is what the kernel enforces on the cgroup. Set requests too low and your pod lands on an already-full node; set limits too low and you meet the OOM killer. The two numbers answer different questions, and conflating them is how clusters get "mysterious" performance problems.
Read before you set
kubectl top pod --containers kubectl describe node <node> | grep -A8 "Allocated resources"
Collect a week of real usage. Your request should sit near the 90th percentile of observed usage, not the average — the average is what you exceed half the time.
Three patterns that work
- Stateless web services: requests = p90 CPU, limit = 2× request. Memory: request = peak + 20%, limit = request (do not let memory burst — it gets killed, not throttled).
- Batch/CI jobs: high CPU limits (or no CPU limit), but always a memory limit. CPU throttling slows a batch job harmlessly; memory overage kills it.
- Databases and stateful sets: requests == limits (Guaranteed QoS) so the node never evicts them first, and the kernel never surprises them.
The anti-patterns
CPU limits on latency-sensitive services cause silent throttling artifacts (p99 spikes at random). Memory requests == tiny + huge limit causes node overcommit and eviction storms. And never copy the chart's defaults — they are placeholders, not recommendations.
Want a sandbox to break these rules on purpose? A managed cluster costs cents per node-hour: DOKS on DigitalOcean. (Partner link — never costs you extra.)