Postgres "the database system is starting up" — Crash Recovery or Stuck?

After an unclean shutdown Postgres replays WAL before accepting connections. Normally minutes; when it never ends, recovery is stuck — and killing it makes things worse.

What you'll see

Root causes

Normal crash recovery replaying WAL

Postgres is doing its job: replaying WAL to the last consistent point. Huge WAL volumes (long-running bulk loads, slots pinning WAL) take time. Watch the log — the LSN position advances.

Recovery genuinely stuck (rare)

Disk errors, a corrupt segment, or archive/restore commands hanging. The log stops advancing and repeats the same LSN.

Fix it

  1. Watch the startup log — is the LSN advancing?
    sudo journalctl -u postgresql -f | grep -iE 'redo|consistent|record|lsn'   # moving forward = be patient
  2. Estimate the remaining replay
    # compare 'redo starts at' vs current in log; or pg_wal size on disk shrinking over time = progress
  3. If stuck on the same position, check for I/O errors or a hung restore_command
    dmesg -T | tail -20   # disk errors; and SHOW restore_command-related config in postgresql.conf if in standby mode
  4. Only after confirmed disk corruption: restore from base backup, not by poking data files
    # pg_basebackup/PITR from your latest backup; never 'fix' a corrupt cluster by deleting pg_wal segments

Field note

Restarting the server mid-recovery restarts replay from scratch (well, from the last checkpoint) — it feels productive and lengthens the outage. Prevention: checkpoint_timeout and max_wal_size tuned for your write load keep recovery windows short.

Common questions

How long should crash recovery take?

Roughly proportional to the WAL since the last checkpoint — often seconds to a few minutes. If the log's redo position advances, wait. If it doesn't move for many minutes, you have a real I/O or restore problem.

Can I connect during recovery to check progress?

Not normally — but the server log (journalctl/docker logs) narrates progress. In containers, docker logs -f shows the same replay messages.

Ship it right the first time

Our most-documented failures, packaged as ready-to-ship starter kits: Docker, Kubernetes, and Terraform.

Browse the template store →

One-time. Yours to modify. Instant download from the NinjaOps template store.