← Back to home

Insight: Infrastructure and Continuity

The point that looks most stable is usually the one that breaks the business.

Infrastructure is not simply what sits in the building. It is an approach that applies wherever the systems actually run: a cloud platform, a server room, or both. The underlying architecture still has to be understood, configured deliberately, and designed so the business keeps operating when a component fails. Because eventually one will.

Illustrative rather than tied to a named client. The principles are identical across Microsoft Azure, Amazon Web Services, Google Cloud and on-premises equipment; only the controls differ.

The Principle

Wherever it runs, the questions are the same

Moving to a cloud platform does not remove the need for design. It changes which controls are available, not whether they need to be considered.

Question One

Is it backed up?

A virtual machine running in a cloud platform is not backed up because it is in the cloud. It is backed up because someone configured that, verified it works, and placed a copy somewhere separate from the original.

Question Two

Is there redundancy?

If any single component stopped right now, whether a connection, a firewall, a switch or a host, would the business continue? For most organisations the honest answer is unknown, because it has never been tested.

Question Three

Is access controlled?

A firewall protecting cloud services, site-to-site connectivity linking locations properly, and services reachable only by those who need them. Availability without control is only half the design.

The question behind all three

What happens if the server room has a fire? It sounds remote until you follow it through: the equipment, the local backup sitting beside it, and the copy nobody moved offsite. A great many businesses would lose all three in the same afternoon. Continuity planning is simply the discipline of asking that question about every part of the estate before circumstances ask it on your behalf.

The Review

Finding the single points of failure

Every environment has them. The work is identifying which ones exist, deciding which are acceptable, and removing the rest. Select any component to see what happens when it fails and what removes the exposure.

Connectivity and power: the dependencies beneath everything

Systems and data: what the business actually runs on

Downtime is not an inconvenience. It is an invoice.

Staff unable to work, orders unable to process, customers unable to reach you. The cost is real whether or not anyone calculates it.

The Decision

How much resilience does the business actually need?

Resilience is a spectrum, and each step upward costs more than the last. The right position depends on what an hour of downtime genuinely costs this organisation.

Level one: recoverable

Backups exist, are held separately from the systems they protect, and have been tested by actually restoring something. The business can recover, but it will lose hours or days doing so. Adequate where that is genuinely tolerable.

Level two: redundant

No single component takes the business down. A second connection on a different technology, paired firewalls, stacked switching. Something failing becomes an event to attend to rather than an outage to announce.

Level three: highly available

Systems continue through a failure with little or no interruption. Appropriate for high-intensity operations where even brief unavailability carries real cost, and priced accordingly.

Level four: continuity planned

The business keeps operating when an entire site is unavailable. Documented, rehearsed and owned, covering not only systems but where people work and how customers are told. Necessary where downtime simply cannot be absorbed.

The calculation worth doing once

Take the number of people unable to work during an outage, multiply by an hour of their cost, add the orders not processed and the customers who could not reach you. Set that figure against what removing the exposure would cost. In most organisations the answer is immediate and uncomfortable, and it converts a technical recommendation into a business decision that can be signed off in a board paper.

The Point

Better safe than sorry, costed properly

A great many businesses run on single points of failure and single points of backup, and they do so comfortably, because those components have been reliable for years. Reliability is not resilience. The component that has never failed is simply one that has not failed yet, and its long record of stability is often exactly why nobody has thought about what happens when it does.

I have seen that point, the stable one, the one nobody worried about, become the thing that stopped a business. Investment in continuity is not spending against a problem you have. It is spending against the one you will eventually have, at a price you choose rather than one chosen for you.