MySQL: "Server Has Gone Away" (Error 2006) — The Four Real Causes

The server closed the connection (or it never survived). Timeout expiry, packet overflow, server restart, or network killing idle TCP. The error's timing — during idle vs during a big query — identifies which.

What you'll see

Root causes

wait_timeout / interactive_timeout expiry

The server closes idle connections after wait_timeout (default 8h, often tuned to minutes). Apps that hold a connection open between requests without validating find it dead. Error 2006 on first query after a pause is the signature.

max_allowed_packet overflow

A single statement bigger than max_allowed_packet drops the connection (not a clean error). Imports/JSON blobs/batch INSERTs mid-query = this cause. Server-side variable vs client max_allowed_packet on the driver both matter.

Server restart or kill / network idle-kill

mysqld restarted (check uptime), or a NAT/firewall silently drops idle TCP — the classic cloud-database middlebox problem: your side sees ESTABLISHED, the other side closed long ago.

Fix it

  1. Distinguish by WHEN it fails
    # after idle -> timeouts; mid-big-query -> packet size; random + load -> server/network
  2. Timeout expiry: validate/reconnect instead of raising limits
    SHOW VARIABLES LIKE 'wait_timeout';   # app side: connection pool ping/validate-on-checkout
  3. Packet overflow: raise on BOTH ends
    SET GLOBAL max_allowed_packet=256*1024*1024;   # + client mysql.cnf / driver option (they negotiate the min)
  4. Middleman idle drops: keepalives
    # my.cnf [mysqld]: interactive_timeout / and set TCP keepalive: net.ipv4.tcp_keepalive_time=300 on the app host

Field note

Error 2013 (lost connection during query) is the same family mid-stream; 2006 is the stale-handle variant. Both trace to the same four causes. Polling SHOW GLOBAL STATUS LIKE 'Aborted_clients' before and after a failure window tells you whether the server or the network is the closer.

Common questions

Why does it work fine under local dev but fail in production?

Production adds middleboxes (NAT, load balancers, cloud network policies) with shorter idle timeouts than your dev loop, plus a longer-lived app process that actually hits wait_timeout. Same code, different environment lifetimes.

Should I just retry on 2006?

Retry with a FRESH connection (never on the dead handle) is correct resilience for the timeout/idle class — but fix the root cause too: validate-on-checkout in the pool eliminates most of it.

Ship it right the first time

Our most-documented failures, packaged as ready-to-ship starter kits: Docker, Kubernetes, and Terraform.

Browse the template store →

One-time. Yours to modify. Instant download from the NinjaOps template store.