Kubernetes Pods Can't Resolve Service Names (CoreDNS/ndots Debug)

nslookup works for external names but service-name lookups fail — or everything DNS is flaky. Two usual suspects: CoreDNS being down, and the ndots:5 search-domain behavior making every lookup slow or wrong.

What you'll see

Root causes

CoreDNS down, misconfigured, or under-provisioned

kubectl -n kube-system get pods -l k8s-app=kube-dns — check ready state and logs. The corefile (forward, rewrite rules) can also point at a dead upstream.

ndots:5 search path explosion

Cluster DNS config makes every non-FQDN lookup try up to 6 suffixes first. High-concurrency pods multiply CoreDNS queries by 6 — the classic cause of intermittent DNS under load. FQDN with a trailing dot or fewer ndots fixes it.

DNS-based service discovery expectations mismatch

The name format matters: <svc>.<namespace>.svc.cluster.local — cross-namespace calls need the namespace part; same-namespace can use the bare name.

Fix it

  1. Test DNS from inside a pod, both forms
    kubectl run dnstest --rm -it --image=busybox --restart=Never -- nslookup kubernetes.default.svc.cluster.local
  2. Check CoreDNS health and config
    kubectl -n kube-system get pods -l k8s-app=kube-dns -o wide && kubectl -n kube-system logs deploy/coredns --tail=20
  3. Mitigate ndots amplification in your workloads
    # use FQDNs with trailing dot: curl https://api.example.com.  ; or dnsConfig: options: [{ name: ndots, value: "2" }] on pods hitting external APIs hard
  4. Scale CoreDNS if it's genuinely saturated
    kubectl -n kube-system scale deploy/coredns --replicas=4 ; # watch: kubectl -n kube-system top pods -l k8s-app=kube-dns

Field note

dnstest pod disappears after exit (--rm) — safe diagnostic anywhere. headless services and ExternalName have their own resolution quirks — check the service type when only one service misresolves.

Common questions

What is ndots and why does it slow my pods down?

resolv.conf search logic: each lookup with fewer than 5 dots tries every search suffix first. So 'api.example.com' (2 dots) triggers multiple failed cluster lookups before the real one. A trailing dot (FQDN) or dnsConfig ndots=2 cuts the query volume hard.

External DNS works but cluster service names don't resolve. What breaks?

Usually CoreDNS or the kube-dns service itself (check the coredns pod and its readiness) — or the query misses the cluster.local suffix. Test from a busybox pod with the full service FQDN to isolate which.

Ship it right the first time

Kustomize base with probes, PDBs, and zero-downtime rollouts already wired.

Kubernetes Production Blueprints — $27 →

One-time. Yours to modify. Instant download from the NinjaOps template store.