A 429 is an application-level 'slow down', not a network failure. The response headers usually tell you exactly when to retry and how — clients that honor them fix themselves.
Rate limits are per token, per IP, or per endpoint. Check the vendor's docs against your request pattern — fan-out loops and retry storms are the usual offenders.
The 429 usually carries Retry-After or X-RateLimit-Reset headers. Clients that ignore them and hammer anyway get deprioritized further or temporarily blocked.
curl -sI https://api.example.com/v1/things | grep -iE 'retry-after|rate-limit|429'
# retry on 429/503: wait = min(base * 2^n, cap) + random_jitter; honor Retry-After when present
# e.g. p-limit in Node, ratelimit in Python, or a token bucket in front of fan-out calls
# your own nginx? limit_req settings; your own app? middleware defaults — 429s from your stack are yours to tune
Shared pools: some providers count by IP for anonymous traffic — a NAT office can trip limits for everyone. Log 429s with the retry window; alerting on them without the window tells you nothing actionable.
Never — immediate retries add load and often extend throttling. Honor Retry-After if present, otherwise exponential backoff with jitter, and cap total retries.
If it's a third-party API: higher-tier plans, API keys (per-token limits usually beat per-IP), or asking for a quota bump. If it's your own service: tune the limiter middleware deliberately.
Our most-documented failures, packaged as ready-to-ship starter kits: Docker, Kubernetes, and Terraform.
Browse the template store →One-time. Yours to modify. Instant download from the NinjaOps template store.