After an unclean shutdown Postgres replays WAL before accepting connections. Normally minutes; when it never ends, recovery is stuck — and killing it makes things worse.
Postgres is doing its job: replaying WAL to the last consistent point. Huge WAL volumes (long-running bulk loads, slots pinning WAL) take time. Watch the log — the LSN position advances.
Disk errors, a corrupt segment, or archive/restore commands hanging. The log stops advancing and repeats the same LSN.
sudo journalctl -u postgresql -f | grep -iE 'redo|consistent|record|lsn' # moving forward = be patient
# compare 'redo starts at' vs current in log; or pg_wal size on disk shrinking over time = progress
dmesg -T | tail -20 # disk errors; and SHOW restore_command-related config in postgresql.conf if in standby mode
# pg_basebackup/PITR from your latest backup; never 'fix' a corrupt cluster by deleting pg_wal segments
Restarting the server mid-recovery restarts replay from scratch (well, from the last checkpoint) — it feels productive and lengthens the outage. Prevention: checkpoint_timeout and max_wal_size tuned for your write load keep recovery windows short.
Roughly proportional to the WAL since the last checkpoint — often seconds to a few minutes. If the log's redo position advances, wait. If it doesn't move for many minutes, you have a real I/O or restore problem.
Not normally — but the server log (journalctl/docker logs) narrates progress. In containers, docker logs -f shows the same replay messages.
Our most-documented failures, packaged as ready-to-ship starter kits: Docker, Kubernetes, and Terraform.
Browse the template store →One-time. Yours to modify. Instant download from the NinjaOps template store.