Disaster recovery testing for companies
A disaster recovery test examines whether the backups you hold are enough to rebuild central company systems outside your own infrastructure. In the ITN Recovery Lab that happens on separate infrastructure — your production environment is left untouched.
What is settled before the test
The quality of a disaster recovery test is decided in the preparation. Five points, without which a test leaves you with nothing but a feeling.
Scope
Which systems are part of the test? What makes sense is the core that carries the business — not everything that exists.
Order
Which system has to be running first? Directory service and name resolution before the database, database before the application server. In an emergency, that order is exactly the information nobody has.
Known dependencies
What do you already know about services waiting on each other, hard-coded addresses and licence servers? The test finds the rest.
Success criteria
Which processes have to be demonstrably working at the end? You set those criteria — and you check them yourself later.
Roles
ITN provides the isolated environment, carries out the restore and starts the systems. You check that things work in business terms and sign off. On request we document the result.
- ITN: restore and boot
- Customer: functional testing
- Documentation on request
Check your targets in practice instead of merely defining them
In many companies RTO and RPO sit in a document: recovery should take four hours, the maximum acceptable data loss is one hour. Both figures are requirements — not measurements. Whether they can be met is something nobody knows until somebody has tried.
A test moves the discussion from assumption to observation. It shows how long restoring the chosen systems actually took, where it stalled, and what state the data was in at the end. Out of that comes either confirmation of your targets — or a well-founded correction to them.
Worth being clear about: the time measured in a test is not a commitment for a real emergency. It is a solid indication of where you stand — more than a figure in a concept paper, less than a guarantee.
The time needed comes out of the test, not out of a vendor's figure.
Usually it is not the volume of data that holds the restart up, but a dependency.
You can see on the restored system what state the data is actually in.
A part of the test that fails is not a poor result — it is the actual value of doing it.
What a disaster recovery test is not
Three clear limits, so that the expectation is right.
Not protection
A test prevents neither an outage nor an attack. It only tells you whether you can start again afterwards.
Not forensics
We do not investigate attacks and do not analyse traces. A test is a planned exercise, not emergency response.
A record
The result is a checked statement about your ability to restart, as at the date of the test — documented on request.
Frequently asked questions about disaster recovery testing
What belongs in a disaster recovery test plan?
The scope of the systems to be tested, the order in which they are restored, the known dependencies, the success criteria, and the roles: who restores, who checks in business terms, who signs off. Without agreed success criteria you cannot say afterwards whether the test succeeded.
How do RTO and RPO relate to a test?
RTO and RPO are targets a company sets for itself. A test shows whether those figures are achievable under real conditions. Only once a restore has actually been carried out can the time it takes be assessed rather than guessed.
Does ITN commit to a particular recovery time?
No. We do not promise a recovery time in advance, because it depends on your environment — the volume of data, the number of systems, the dependencies and the format of the backups. What the test actually took is recorded in the documentation on request. That is a measurement, not a commitment.
How does this differ from an emergency handbook?
An emergency handbook describes what should happen when things go wrong. A test shows whether it works. The two belong together: what a test turns up is usually the most valuable correction to the handbook, because it comes from practice rather than from assumption.
Does the test have to cover the whole of our IT?
No, and usually that would not be sensible either. What tells you something is the core that carries the business. A test over a few well-chosen systems says more than a sweep across everything with no clear success criteria.
More on this: cyber recovery, for when your own infrastructure can no longer be trusted, or testing backup restores as a starting point. Overview in the Recovery Lab.
