High load means runnable tasks are waiting. It can be CPU, I/O, or both — and treating the wrong one makes it worse. A five-command triage tells you which resource is the real bottleneck.
load / cores > 1 sustained with high %us/%sy. Find the consumer: top (P to sort by CPU), pidstat -u 1, then drill into threads (top -H -p <pid>).
High %wa in top means processes sleep on disk/NFS. iostat -x 1 shows util and await per device; pidstat -d 1 shows who does the I/O. Slow disks masquerade as 'app slowness'.
ps -eo state,pid,cmd | grep '^D' — processes stuck in disk/network I/O inflate load while consuming no CPU. NFS hangs and dying disks are the classics.
nproc && uptime && top -bn1 | head -5 # read the %Cpu(s) line: us vs wa vs sy vs st
iostat -x 1 3 2>/dev/null | tail -20; pidstat -d 1 3 2>/dev/null | tail -10 # apt install sysstat if missing
ps -eo pid,pcpu,pmem,comm --sort=-pcpu | head -8; top -H -p <pid> # strace -c -p <tid> for syscall-level digging
# %st = steal time. That's the hypervisor servicing other tenants — nothing to fix in-guest; resize or migrate the VM
ps -eo state,pid,wchan:20,cmd | awk '$1=="D"' | head
Load average counts running + uninterruptible tasks — that's why it can be huge with idle CPUs during disk hangs. vmstat 1 is the fastest single overview: r (runnable), b (blocked), wa (iowait), si/so (swap).
Compare to core count: load 4 on 8 cores is fine; load 8 sustained is saturated (CPU or I/O). The number alone doesn't say which — the %Cpu wa/us split and D-state check do.
It's I/O: tasks blocked on disk or NFS (D-state) count toward load. Run iostat -x 1 and look for high util/await on a device, then find the writing process with pidstat -d.
Our most-documented failures, packaged as ready-to-ship starter kits: Docker, Kubernetes, and Terraform.
Browse the template store →One-time. Yours to modify. Instant download from the NinjaOps template store.