Back to blog

The Backup Isn't the Plan. The Restore Is.

A data center fire in New Delhi is the same lesson infrastructure keeps re-teaching: a backup nobody has restored under pressure is not disaster recovery. It's theater.

disaster recoveryinfrastructureoperationsreliability

The Backup Isn't the Plan. The Restore Is.

A data center fire isn't a black swan. Data centers burn. It happened again in New Delhi.

A fire at the STT Global Data Centres India facility, owned by ST Telemedia and Tata Communications, left customers fearing the loss of decades of data and disrupted Google Cloud network traffic across India. One customer, Matrix Cellular, said it may have lost access to more than 20 years of operational and business data.

The lesson isn't that data centers can burn. Anyone who has run infrastructure long enough already knows that. The lesson is simpler and more brutal:

A backup you have not restored under pressure is not a backup. It is a hope with a purchase order.

Google's own service-health page said the fire required an emergency power shutdown of networking equipment, isolated a Delhi point of presence, and reduced available network capacity. Customers in Delhi, Chennai, Mumbai, and surrounding areas saw intermittent latency and possible packet loss. On June 23, Google still listed "no workaround" and said customers could see elevated latency until the affected facility was fully restored.

That phrase -- no workaround -- is where operators should sit up straight.

Most companies think they have continuity because they bought something with "redundant," "geo," "replicated," or "managed" in the name. Tape libraries, SAN replication, DR colo, VMware clusters, Kubernetes, cloud regions, SaaS exports, now AI workflows. The label changes. The failure mode does not.

The system works right up until someone asks the only question that matters: Can we restore the business, not just the data?

Tata Communications said it had restored services for customers who subscribed to recovery and backup services, while further efforts continued "to the extent possible." In a failure, your vendor only owes you what the contract says -- nothing more boots a server, rebuilds an index, replays transactions, restores identity, reconnects DNS, or gets billing running again.

This is where operators need to stop accepting noun-based answers. "Do we have backups?" is a weak question. Ask instead:

"When was the last full restore test?" "How long did it take?" "What data was missing?" "Who signed the evidence?" "What dependencies failed during restore?" "What happens if the building, provider account, identity plane, and network path are all impaired at the same time?"

That last one is the one most plans fail. The backup exists, but the admin credentials are in the same tenant. The restore guide is in the same wiki. The network route assumes the same facility. The export is encrypted with a key stored in the same cloud account. The vendor can recover the storage, but not your application state. Congratulations, you own a very expensive coffin with blinking lights.

Resilience isn't a feature you buy. It's a habit you prove with a receipt -- date, scope, duration, gaps, owner, next fix, not a policy binder, a dashboard, or a happy-path screenshot from 18 months ago. If the restore has never been run cold, by someone who did not build it, from documentation that lives outside the blast radius, then you do not have disaster recovery. You have theater.