Cncf iconCncfSep 10, 2026 ~7 min source read

Kubernetes disaster recovery: what backups alone don’t solve

A CNCF lab shows three reproducible failure scenarios that separate having backups from being able to recover. The practical gaps are at the joins between recovery layers: confirming bytes moved, separating declared state from stored state, and preparing the environment the backups are restored into.

Kubernetes disaster recovery: Guidance from three reproducible failure scenarios

Share this story

Send the public story page.

Useful takeaways from this story.

The document covers recovery of stateful applications running on Kubernetes: verifying that backups contain data, the split between declared state and stored state, and consistency across multi-volume...

It does not cover compliance frameworks, product comparisons, or recovery of the underlying cloud or datacenter infrastructure, though it names where those responsibilities begin.

Specific tools appear where a scenario needs them (Velero, the CSI snapshot APIs).

This CNCF lab walks through failure scenarios for stateful applications on Kubernetes to show why having backups is necessary but not sufficient for recovery. Each scenario is reproducible locally and uses concrete tooling (Velero, a data mover, the CSI snapshot APIs where needed) as reference implementations. The goal is to separate two different things teams must recover: declared state (YAML, Kubernetes objects) and stored state (persistent volume data).

Four recovery layers and where recovery breaks

The document frames recovery as four layers that all must come back together. Each layer typically has mature tooling, but recovery commonly fails at the joins between layers. Examples of join failures covered by the lab include:

  • A restored cluster with no data because volume bytes were never verified as moved.
  • Restored data with no traffic path because networking or load balancers weren't recreated.
  • An application definition that provisions an empty volume because Git only contained declarations.

The lab uses two preexisting local clusters (production and recovery), an S3-compatible object store outside both clusters for backups, a Git service for manifests, and a small PostgreSQL workload with known contents (four rows) so restores can be validated.

Scenario 1 — verifying that a backup contains data

Backup tools protect two parts: resource definitions and persistent volume data. Volume protection can be via CSI snapshots, provider snapshots, file system backup, or moving snapshot bytes to an external store. In the lab, Velero plus a data mover is used to copy volume bytes into the external S3-compatible store.

A common operational mistake is trusting the backup's Completed status alone. The lab shows how to query the data mover for actual bytes transferred (for example, a datauploads status column reporting bytesDone) to confirm that volume data left the cluster and landed in the external store. That numeric confirmation is a stronger signal than a Completed backup phase.

Even when bytes moved are confirmed, three needed follow-ups remain:

  • Database/application consistency: backups don't automatically quiesce databases. Flush or quiesce hooks must be configured when the application requires them.
  • Storage transformations: restoring onto different infrastructure may require storageClass mappings or other transformations.
  • End-to-end testing: only a full recovery test proves the application starts, contains expected data, and serves traffic.

Scenario 2 — declared state is not stored state (the GitOps trap)

In this scenario production is powered off and the recovery cluster, which already exists, runs a GitOps controller pointed at Git and a backup tool pointed at the shared store. The Git repository contains only the application manifests. The GitOps controller successfully applies the manifests: StatefulSet rolls out, pods are Running and Ready, and dashboards look healthy.

Practical takeaway: You need both Git (or other declared-state tooling) and a backup/restore process for persistent volumes. A DR plan that begins with "restore the backup" must also state what the backup will be restored into and how storage mappings, cluster infrastructure, and application consistency will be handled.

Boundary: restoring isn't the same as recreating infrastructure

Backup tools restore objects into an existing cluster. They do not create the cluster control plane, nodes, networking, load balancers, or DNS. Recovering Kubernetes itself belongs to infrastructure-as-code, Cluster API, or other platform-level tooling. Teams must define which system is responsible for bringing up the environment that receives restored resources.

Design recovery as coordinated steps across layers, verify that volume bytes are present in the backup store, test end-to-end restores into a recovery cluster, include storage mapping and quiesce hooks for stateful apps, and document exactly what environment backups are restored into.

More context around this story.

[Перевод] Аварийное восстановление в Kubernetes: три воспроизводимых сценария сбоя
Habr iconHabrSep 24, 2026

[Перевод] Аварийное восстановление в Kubernetes: три воспроизводимых сценария сбоя

Команда VK Cloud перевела статью, описывающую три сценария сбоя, которые отличают наличие резервных копий от возможности восстановления, а также рекомендации, следующие из каждого сценария. Каждый сценарий воспроизводим на ноутбуке из указанного выше репозитория стенда, и каждый показанный вывод терминала является реал

Best practices for handling cloud reliability incidents
Google iconGoogleSep 15, 2026

Best practices for handling cloud reliability incidents

Cloud outages can range from global service disruptions to issues isolated to a specific region, zone, or even just your project, workload or application. If you suspect a Google Cloud Platform outage is impacting your services, we recommend you follow a structured “Verify→ Investigate→Report→Resolve→Review" workflow t

Loading more related stories...

Keep reading in the app

Open the app view to save this story, compare related coverage, and continue from the same source.

Open in app