My home lab is a seven-node kubeadm cluster that runs real services for my family and my own projects. In August it hit a wall that many production clusters are about to hit too, so I used it as a rehearsal.
Three problems that blocked each other
- Kubernetes 1.33 went end-of-life in June 2026.
- The community ingress-nginx was retired (archived in March 2026, no more security patches) and its last release did not support anything newer than 1.33. So it blocked every upgrade hop.
- The original kubespray inventory was gone, so nothing could be re-applied declaratively. The cluster had quietly become a plain kubeadm cluster.
You cannot fix these one at a time: the ingress blocks the upgrade, and the upgrade is the reason to change the ingress. So the plan became one day: Calico → Cilium (eBPF, kube-proxy replaced), ingress-nginx → Gateway API, and 1.33 → 1.34 → 1.35 → 1.36.
Why Cilium, knowing it was the riskier option
The safer, CNI-agnostic choice was a standalone Gateway API implementation such as Envoy Gateway. I chose Cilium anyway and wrote down why: Gateway API resources are portable. The HTTPRoutes written today survive a later change of GatewayClass. The riskier choice is not a dead end, unlike bumping the retired ingress controller one more time, which would still have stopped at 1.35.
Writing the reasoning down paid off later: when the same fork came up on a work cluster, the answer was already there.
Live migration: tried, and abandoned on evidence
Cilium documents a per-node migration where Calico and Cilium run side by side. It failed for a concrete reason: Calico ran VXLAN in Always mode, and Calico nodes had no route to Cilium’s pod CIDR. A migrated node’s pods could reach nothing on unmigrated nodes — not even CoreDNS. That is two islands, exactly what the hybrid mode is supposed to prevent, and upstream warns the procedure is untested with existing NetworkPolicy providers.
So I switched to a bounded big-bang: label all nodes, write the Cilium CNI config, reboot the whole cluster (about 2.5 minutes of downtime), then reboot once more after removing Calico to clear its iptables rules and VXLAN interface.
The order that mattered
- kube-proxy replacement had to land together with removing NodeLocal DNSCache and pointing
clusterDNSback at CoreDNS. Once Cilium translates ClusterIPs in eBPF, the node-local cache can no longer work, and pods would have been left pointing at a dead resolver. - Cilium plus kube-proxy in IPVS mode is a broken middle state, not a waypoint. Pods reach no ClusterIP at all.
- The ingress cutover was the low-risk part. The Gateway ran next to ingress-nginx, every hostname was verified through it with a
Host:header while real traffic still went the old way, then hosts were switched one by one with instant rollback. The Cloudflare tunnels, whose rules live in the Cloudflare dashboard rather than in the cluster, were repointed through the API. Only then was ingress-nginx deleted.
Traps worth reading before you try this
- The chart’s default pod CIDR is
10.0.0.0/8. On a network with private routes in10.xit silently swallows your LAN and every VPN route. - If
/opt/cni/binis owned by a non-root user, Cilium’s init container fails: it dropsDAC_OVERRIDE. - With kube-proxy gone, the agent needs
k8sServiceHostpointing at the local API endpoint, or it deadlocks on its own eBPF datapath. helm upgradedoes not restart Cilium when only the ConfigMap changed. The agent and the operator both need arollout restart. The built-in drift check tells you (Mismatch found).TLSRoutelives in the experimental Gateway API CRD set.- A proxy that sends no SNI breaks TLS passthrough with a 502.
kubeadm upgrade nodeon the last control-plane node re-applied the kube-proxy addon on every hop, because--skip-phasesis not inherited. Check for kube-proxy after every hop.
Result
Seven of seven nodes Ready on 1.36, zero problem pods, every storage volume healthy, every Argo CD application Synced and Healthy, every endpoint answering. The lesson I carry into work clusters is not “big-bang is fine” but “measure first, then pick the migration shape”: the live migration looked safer on paper and was the one that would have broken everything.