Nginx returns 503 when worker_connections (default 512/1024) can't cover concurrent sockets. Raising it is one line — but check ulimit and what's actually holding connections first.
Each worker counts every connection (client + proxied upstream pairs use two). worker_connections x worker_processes is the theoretical max — defaults are tiny for busy proxies.
worker_connections can't exceed the worker's ulimit (often 1024 from systemd defaults). Raising nginx's number without LimitNOFILE does nothing. worker_rlimit_nofile sets it per-worker.
Every slow backend request parks a connection pair. The real fix may be upstream timeouts (proxy_read_timeout) instead of more slots.
sudo tail -50 /var/log/nginx/error.log | grep -E 'worker_connections|1024|open files'
# nginx.conf events block: worker_connections 4096; plus (top level): worker_rlimit_nofile 8192;
# nginx.service [Service]: LimitNOFILE=8192 ; then: sudo systemctl daemon-reload && sudo systemctl restart nginx
# proxy_read_timeout 30s; proxy_connect_timeout 5s; — measure first: nginx stub_status (active connections) tells you the real concurrency
Add the http_stub_status_module endpoint (localhost-only) to watch 'active connections' over time — data before knob-turning. Keep-alive from browsers means 'active' includes idle-but-open sockets; worker_connections counts them too.
Common production: 4096-16384, paired with an equal-or-higher file descriptor limit. The right number depends on measured active connections, not vibes — watch stub_status before and after.
Two usual suspects: the OS fd limit is still the ceiling (worker_rlimit_nofile + systemd LimitNOFILE), or upstream slowness is parking connections. The error log wording tells you which limit actually refused.
Our most-documented failures, packaged as ready-to-ship starter kits: Docker, Kubernetes, and Terraform.
Browse the template store →One-time. Yours to modify. Instant download from the NinjaOps template store.