ChallengekubernetesDestructive
Node Drain, Upgrade & Recovery — Challenge
Take a node out of service without taking the application with it, and find out which workloads were never ready for it.
- Time
- 27 min
- Level
- Advanced
- Objectives
- 4 objectives
- Cost
- Free
Before you start
You will need
- kind (multi-node)
- kubectl 1.28+
You will be able to
- Cordon and drain a node safely
- Protect availability during voluntary disruption with a PDB
- Recognise workloads that cannot survive rescheduling
Cost — Free
— a multi-node kind cluster. See the setup note; a single node cannot demonstrate rescheduling.
You are done when
0 of 4
The goal#
Achieve the same outcome as Node Drain, Upgrade & Recovery, from an empty starting point, without the steps.
The cluster needs a Kubernetes upgrade. That means taking each node out of service in turn, and the first one you try teaches you which of your workloads were only ever running by luck.
What must be true when you are done#
- A node is drained with no failed requests to the application.
- A PodDisruptionBudget blocks a drain that would breach availability.
- You identified at least one workload that could not be rescheduled cleanly, and why.
- The node returns to service and receives Pods again.
Rules#
- Do not open the guided lab until you are finished, or until the same problem has held you up for 20 minutes.
- Documentation is allowed and encouraged.
- Verify every criterion with a command whose output you can read.
If you get stuck#
- What did you expect, exactly?
- What happened instead — the error text, not a paraphrase?
- Which layer is that error from?
- What is the smallest command that proves the layer below is fine?
The concept behind it
Next up
Lab 57 of 58 on the project path