Two transactions each hold locks the other needs; Postgres detects the cycle and kills one (40P01). Consistent lock ordering in code removes the class; a retry removes the symptom. Both have their place.
Transaction A updates rows 1→2 while B updates 2→1. Postgres detects the cycle and aborts the victim. The fix: always acquire row locks in a consistent order (e.g. ORDER BY id in the UPDATE's inner query).
Some sessions taking FK/share locks, others exclusive, on overlapping sets: foreign keys create hidden row locks that enter cycles even when app code looks conflict-free.
# server log: 'Process 123 waits for ... blocked by Process 456' + the two SQL statements. That pair is the whole diagnosis.
# UPDATE ... WHERE id IN (SELECT id FROM t WHERE ... ORDER BY id FOR UPDATE) -- sorted locking serializes cleanly
# long transactions hold locks longer; commit promptly, move reads and API calls outside the transaction
# on 40P01: rollback, small backoff, retry the transaction — standard practice for the residual race
40P01 rolls back the VICTIM transaction only (Postgres picks the one doing less work); the other proceeds and commits. The victim's work is lost but the database stays consistent — retry is safe by design. deadlock_timeout (default 1s) is detection latency, not lock-wait duration: lowering it finds cycles faster at higher detection overhead. It's rarely the knob you're looking for.
No — that's the point of detection: one transaction is cleanly rolled back, the other commits, the database never applies the inconsistent interleaving. The victim's client sees an error and can retry.
Lock cycles need concurrent timing: sequential test runs can't interleave. Reproduce with concurrent clients hitting the same code path, then fix the ordering — the pair of statements in the log shows you exactly where.
Our most-documented failures, packaged as ready-to-ship starter kits: Docker, Kubernetes, and Terraform.
Browse the template store →One-time. Yours to modify. Instant download from the NinjaOps template store.