This CNCF lab walks through failure scenarios for stateful applications on Kubernetes to show why having backups is necessary but not sufficient for recovery. Each scenario is reproducible locally and uses concrete tooling (Velero, a data mover, the CSI snapshot APIs where needed) as reference implementations. The goal is to separate two different things teams must recover: declared state (YAML, Kubernetes objects) and stored state (persistent volume data).
Four recovery layers and where recovery breaks
The document frames recovery as four layers that all must come back together. Each layer typically has mature tooling, but recovery commonly fails at the joins between layers. Examples of join failures covered by the lab include:
- A restored cluster with no data because volume bytes were never verified as moved.
- Restored data with no traffic path because networking or load balancers weren't recreated.
- An application definition that provisions an empty volume because Git only contained declarations.
The lab uses two preexisting local clusters (production and recovery), an S3-compatible object store outside both clusters for backups, a Git service for manifests, and a small PostgreSQL workload with known contents (four rows) so restores can be validated.
Scenario 1 — verifying that a backup contains data
Backup tools protect two parts: resource definitions and persistent volume data. Volume protection can be via CSI snapshots, provider snapshots, file system backup, or moving snapshot bytes to an external store. In the lab, Velero plus a data mover is used to copy volume bytes into the external S3-compatible store.
A common operational mistake is trusting the backup's Completed status alone. The lab shows how to query the data mover for actual bytes transferred (for example, a datauploads status column reporting bytesDone) to confirm that volume data left the cluster and landed in the external store. That numeric confirmation is a stronger signal than a Completed backup phase.
Even when bytes moved are confirmed, three needed follow-ups remain:
- Database/application consistency: backups don't automatically quiesce databases. Flush or quiesce hooks must be configured when the application requires them.
- Storage transformations: restoring onto different infrastructure may require storageClass mappings or other transformations.
- End-to-end testing: only a full recovery test proves the application starts, contains expected data, and serves traffic.
Scenario 2 — declared state is not stored state (the GitOps trap)
In this scenario production is powered off and the recovery cluster, which already exists, runs a GitOps controller pointed at Git and a backup tool pointed at the shared store. The Git repository contains only the application manifests. The GitOps controller successfully applies the manifests: StatefulSet rolls out, pods are Running and Ready, and dashboards look healthy.
Practical takeaway: You need both Git (or other declared-state tooling) and a backup/restore process for persistent volumes. A DR plan that begins with "restore the backup" must also state what the backup will be restored into and how storage mappings, cluster infrastructure, and application consistency will be handled.
Boundary: restoring isn't the same as recreating infrastructure
Backup tools restore objects into an existing cluster. They do not create the cluster control plane, nodes, networking, load balancers, or DNS. Recovering Kubernetes itself belongs to infrastructure-as-code, Cluster API, or other platform-level tooling. Teams must define which system is responsible for bringing up the environment that receives restored resources.
Design recovery as coordinated steps across layers, verify that volume bytes are present in the backup store, test end-to-end restores into a recovery cluster, include storage mapping and quiesce hooks for stateful apps, and document exactly what environment backups are restored into.