Kubernetes: Weird TLS/Cert Errors Everywhere (Check the Clock)

Expired-looking x509 errors, 'token invalid', or nodes failing to join with intact certs is the classic symptom of clock skew. Certificates are time contracts.

What you'll see

Root causes

NTP daemon not running or unreachable

VMs resuming from pause, hosts with blocked NTP egress, or chrony/ntpd stopped — the clock free-runs and drifts.

Hypervisor snapshot bring-back

Restored VMs wake up in the past or future; skew exceeds the TLS validity window and every cert check fails.

Fix it

  1. Check time sync status on the affected host
    timedatectl   # 'System clock synchronized: yes'?
  2. Check the skew error detail — expired vs not-yet-valid tells you direction
    sudo journalctl -u kubelet | grep -i 'x509' | tail -3
  3. Fix and verify the sync daemon
    sudo systemctl enable --now chronyd && chronyc tracking
  4. Then restart the affected components
    sudo systemctl restart kubelet   # and containerd if needed

Field note

'Certificate is not yet valid' means the host clock is AHEAD; 'has expired' means it's BEHIND. That one word tells you which direction the skew went before you even check chrony.

Common questions

Why do clock skew and certificates interact?

x509 certs are valid only within a time window: a node clocked even a few minutes off makes valid certs look expired (or not-yet-valid). The API server rejects the node's client cert, kubelet fails auth — a cascade from a time sync problem.

How do I fix it permanently?

Sync time (chrony/ntp) on every node — and on VM hosts, since nested clock drift propagates. Check kubectl get nodes for the resulting NotReady after sync: the cert error clears without rotation once clocks agree.

Ship it right the first time

Kustomize base with probes, PDBs, and zero-downtime rollouts already wired.

Kubernetes Production Blueprints — $27 →

One-time. Yours to modify. Instant download from the NinjaOps template store.