
What happens when the failure you never imagined actually happens? At the inaugural AWS Resiliency Day, UK North, this month, Mark Sailes shared the lessons he’s learned behind building resilient systems, preparing teams for disruption and making sure hard-won experience benefits everyone.
Resilience planning tends to focus on the things we know can go wrong. We map dependencies, identify failure points and design systems to withstand disruption. But what happens when the failure that takes you offline is something nobody thought to plan for? In my experience, that's where the real lessons begin.
One of the defining moments in my own resilience journey happened in a former role.
Our data centre was based in Guernsey. A fishing boat dropped its anchor, and it dragged along the seabed severing all three undersea fibre optic cables connecting the island. Our entire data centre was suddenly cut off.
This is a good example of the kind of event that makes resilience planning so challenging. You can plan for component failures, software bugs and network outages, but it's much harder to anticipate every combination of circumstances that could take a critical service offline.
At the time, our disaster recovery arrangements relied on a cold standby on the mainland. It was there for exactly this kind of situation but having a backup environment and knowing that it will work when you need it are two very different things.
After the incident, we spent the following year carrying out fortnightly resilience tests to make sure the failover would actually work. That was an enormous amount of operational effort just to maintain confidence in our ability to recover.
Eventually, our team decided to take a different approach. We rebuilt on AWS using native regional services, embedding resilience into the architecture rather than relying on a cold standby and repeated manual fail overexercises. The result was a more resilient solution, less operational toil and, importantly, more time to improve the system rather than continually rehearse its recovery.
This experience shaped how I think about resilience. You can't anticipate every failure, but you can change how you design systems, how you prepare people and how you learn when something goes wrong.
When an incident happens, there's often one engineer who gets to the bottom of it. They spend hours investigating, find the root cause, implement a fix and get the service back up. Everyone is relieved, the incident is closed and the team moves on.
But what happens to everything that engineer learned at three o'clock in the morning, especially if they ultimately leave the business? Too often, the knowledge stays with the person who fixed the problem. The next engineer encounters the same failure months later and has to work it all out again. We need to make sure that doesn't happen.
One way to do this is through game days, these are structured exercises in which teams deliberately recreate failure scenarios and practice responding to them in a safe environment. After resolving an incident, ask the engineer who fixed it to lead a game day. Give the wider team their own temporary environment in which they can reproduce the failure, investigate the symptoms and understand why the original problem occurred.
The objective isn't simply to teach everyone how to apply the same fix. It's to help them understand the failure mode and learn how to avoid introducing it in the first place.
By practising in an environment where nothing real is at risk, people can ask questions, experiment and build confidence without the pressure of a live customer incident. It's a practical way to turn individual experience into collective capability.
There's another problem with testing resilience; your own team knows its system better than anyone else. That's a strength, but it can also create blind spots.
When you know how something is supposed to work, it's easy to make assumptions about how it might fail. You naturally test the scenarios you've already considered.
So, try inviting people from outside your team to challenge those assumptions.
In my previous roles, I would bring in technical leads and site reliability engineers (SREs) from other teams and ask them to find ways to break the system. Ideally, they'd have enough freedom to explore failure scenarios the team hasn't anticipated, including tests that are not announced in advance.
This is a form of red teaming, deliberately challenging a system to discover weaknesses before they affect customers.
People who haven't built or maintained your system bring a different perspective. They may question decisions your team takes for granted or identify dependencies nobody has considered.
It also gives your on-call engineers the opportunity to practice diagnosing problems and responding to unexpected behaviour. The key is to make these exercises safe, controlled and focused on learning. The aim isn't to catch someone out, it's to uncover weaknesses while there's still time to address them.
Most organisations are good at celebrating commercial success. We share the new customer, the major deal and the successful launch.
But what about the engineer who eliminates a recurring failure? Or the team that reduces recovery time, removes a fragile dependency or prevents a whole category of incidents? Those achievements deserve to be shared, too.
At AWS, the idea of an "ops win" email is a way to give significant operational improvements the same visibility as other business successes.
The important thing is to do it properly. Explain what the problem was, what changed and what measurable improvement followed. Make the information clear enough that another team, even one that wasn't involved, can understand the lesson and apply it.
Sharing operational improvements helps prevent teams from repeatedly solving the same problems in isolation. It also reinforces a culture in which improving reliability and reducing operational toil are recognised as valuable work.
Resilience isn't just about responding well when things break. It's also about recognising and rewarding the work that makes those failures less likely to happen again.
Of course, how we build systems still matters enormously. My experience with the Guernsey data centre reinforced the value of designing resilience into the architecture, rather than treating recovery as something to bolt on afterwards.
Multi-region architectures are one example. They can help reduce dependence on a single region, but they also introduce challenges that need careful consideration.
For systems that accept writes in multiple regions, for example, ensuring that records remain consistent and duplicate writes are handled correctly can add significant complexity.
Developments in AWS services are making some of these challenges easier to address. Amazon DynamoDB Global Tables with multi-region strong consistency (MRSC), alongside improvements in Amazon Time Sync Service, offer capabilities that can support multi-region designs.
The important point is that we should take advantage of the resilience capabilities available in the platforms we use, while understanding the trade-offs and operational implications of our chosen architecture.
Technology can remove entire categories of manual work, but it doesn't remove the need for teams to understand how their systems behave when something goes wrong. The strongest approach combines resilient architecture with people who know how to operate it.
If there's one message I'd take from these experiences, it's that resilience needs to be an ongoing practice, not a document that gets reviewed once a year.
Here are five practical things teams can start doing:
1. Build resilience in from the start. Consider failure scenarios, dependencies and recovery requirements when designing a service, rather than relying on recovery plans to compensate for fragile architecture.
2. Turn incidents into game days. Recreate significant failures in safe, temporary environments so the whole team learns from the experience, not just the person who resolved the incident.
3. Invite outsiders to challenge your assumptions. Ask people from other teams to test your systems and explore failure scenarios you might never have considered.
4. Share and celebrate operational improvements. Document measurable wins and circulate the lessons so other teams can benefit without having to discover the same problems themselves.
5. Use the resilience capabilities your platform provides. Look for opportunities to automate recovery, reduce manual operational effort and simplify complex resilience arrangements.
You don't need to transform everything at once. Start with one critical workload, identify where the biggest risks lie and choose one practical improvement to make. Then test it, learn from it and build from there.
Ultimately, resilience isn't just about having a backup or ticking a box in a disaster recovery plan. It's about creating the conditions in which people can respond effectively when the unexpected happens and ensuring that every incident leaves us better prepared for the next one.
If you'd like to explore how to build resilience into your projects from the outset, or how to prepare your teams for the failures they haven't yet imagined, we’d love to hear from you. Contact the team here.

What happens when the failure you never imagined actually happens? At the inaugural AWS Resiliency Day, UK North, this month, Mark Sailes shared the lessons he’s learned behind building resilient systems, preparing teams for disruption and making sure hard-won experience benefits everyone.
Resilience planning tends to focus on the things we know can go wrong. We map dependencies, identify failure points and design systems to withstand disruption. But what happens when the failure that takes you offline is something nobody thought to plan for? In my experience, that's where the real lessons begin.
One of the defining moments in my own resilience journey happened in a former role.
Our data centre was based in Guernsey. A fishing boat dropped its anchor, and it dragged along the seabed severing all three undersea fibre optic cables connecting the island. Our entire data centre was suddenly cut off.
This is a good example of the kind of event that makes resilience planning so challenging. You can plan for component failures, software bugs and network outages, but it's much harder to anticipate every combination of circumstances that could take a critical service offline.
At the time, our disaster recovery arrangements relied on a cold standby on the mainland. It was there for exactly this kind of situation but having a backup environment and knowing that it will work when you need it are two very different things.
After the incident, we spent the following year carrying out fortnightly resilience tests to make sure the failover would actually work. That was an enormous amount of operational effort just to maintain confidence in our ability to recover.
Eventually, our team decided to take a different approach. We rebuilt on AWS using native regional services, embedding resilience into the architecture rather than relying on a cold standby and repeated manual fail overexercises. The result was a more resilient solution, less operational toil and, importantly, more time to improve the system rather than continually rehearse its recovery.
This experience shaped how I think about resilience. You can't anticipate every failure, but you can change how you design systems, how you prepare people and how you learn when something goes wrong.
When an incident happens, there's often one engineer who gets to the bottom of it. They spend hours investigating, find the root cause, implement a fix and get the service back up. Everyone is relieved, the incident is closed and the team moves on.
But what happens to everything that engineer learned at three o'clock in the morning, especially if they ultimately leave the business? Too often, the knowledge stays with the person who fixed the problem. The next engineer encounters the same failure months later and has to work it all out again. We need to make sure that doesn't happen.
One way to do this is through game days, these are structured exercises in which teams deliberately recreate failure scenarios and practice responding to them in a safe environment. After resolving an incident, ask the engineer who fixed it to lead a game day. Give the wider team their own temporary environment in which they can reproduce the failure, investigate the symptoms and understand why the original problem occurred.
The objective isn't simply to teach everyone how to apply the same fix. It's to help them understand the failure mode and learn how to avoid introducing it in the first place.
By practising in an environment where nothing real is at risk, people can ask questions, experiment and build confidence without the pressure of a live customer incident. It's a practical way to turn individual experience into collective capability.
There's another problem with testing resilience; your own team knows its system better than anyone else. That's a strength, but it can also create blind spots.
When you know how something is supposed to work, it's easy to make assumptions about how it might fail. You naturally test the scenarios you've already considered.
So, try inviting people from outside your team to challenge those assumptions.
In my previous roles, I would bring in technical leads and site reliability engineers (SREs) from other teams and ask them to find ways to break the system. Ideally, they'd have enough freedom to explore failure scenarios the team hasn't anticipated, including tests that are not announced in advance.
This is a form of red teaming, deliberately challenging a system to discover weaknesses before they affect customers.
People who haven't built or maintained your system bring a different perspective. They may question decisions your team takes for granted or identify dependencies nobody has considered.
It also gives your on-call engineers the opportunity to practice diagnosing problems and responding to unexpected behaviour. The key is to make these exercises safe, controlled and focused on learning. The aim isn't to catch someone out, it's to uncover weaknesses while there's still time to address them.
Most organisations are good at celebrating commercial success. We share the new customer, the major deal and the successful launch.
But what about the engineer who eliminates a recurring failure? Or the team that reduces recovery time, removes a fragile dependency or prevents a whole category of incidents? Those achievements deserve to be shared, too.
At AWS, the idea of an "ops win" email is a way to give significant operational improvements the same visibility as other business successes.
The important thing is to do it properly. Explain what the problem was, what changed and what measurable improvement followed. Make the information clear enough that another team, even one that wasn't involved, can understand the lesson and apply it.
Sharing operational improvements helps prevent teams from repeatedly solving the same problems in isolation. It also reinforces a culture in which improving reliability and reducing operational toil are recognised as valuable work.
Resilience isn't just about responding well when things break. It's also about recognising and rewarding the work that makes those failures less likely to happen again.
Of course, how we build systems still matters enormously. My experience with the Guernsey data centre reinforced the value of designing resilience into the architecture, rather than treating recovery as something to bolt on afterwards.
Multi-region architectures are one example. They can help reduce dependence on a single region, but they also introduce challenges that need careful consideration.
For systems that accept writes in multiple regions, for example, ensuring that records remain consistent and duplicate writes are handled correctly can add significant complexity.
Developments in AWS services are making some of these challenges easier to address. Amazon DynamoDB Global Tables with multi-region strong consistency (MRSC), alongside improvements in Amazon Time Sync Service, offer capabilities that can support multi-region designs.
The important point is that we should take advantage of the resilience capabilities available in the platforms we use, while understanding the trade-offs and operational implications of our chosen architecture.
Technology can remove entire categories of manual work, but it doesn't remove the need for teams to understand how their systems behave when something goes wrong. The strongest approach combines resilient architecture with people who know how to operate it.
If there's one message I'd take from these experiences, it's that resilience needs to be an ongoing practice, not a document that gets reviewed once a year.
Here are five practical things teams can start doing:
1. Build resilience in from the start. Consider failure scenarios, dependencies and recovery requirements when designing a service, rather than relying on recovery plans to compensate for fragile architecture.
2. Turn incidents into game days. Recreate significant failures in safe, temporary environments so the whole team learns from the experience, not just the person who resolved the incident.
3. Invite outsiders to challenge your assumptions. Ask people from other teams to test your systems and explore failure scenarios you might never have considered.
4. Share and celebrate operational improvements. Document measurable wins and circulate the lessons so other teams can benefit without having to discover the same problems themselves.
5. Use the resilience capabilities your platform provides. Look for opportunities to automate recovery, reduce manual operational effort and simplify complex resilience arrangements.
You don't need to transform everything at once. Start with one critical workload, identify where the biggest risks lie and choose one practical improvement to make. Then test it, learn from it and build from there.
Ultimately, resilience isn't just about having a backup or ticking a box in a disaster recovery plan. It's about creating the conditions in which people can respond effectively when the unexpected happens and ensuring that every incident leaves us better prepared for the next one.
If you'd like to explore how to build resilience into your projects from the outset, or how to prepare your teams for the failures they haven't yet imagined, we’d love to hear from you. Contact the team here.

What happens when the failure you never imagined actually happens? At the inaugural AWS Resiliency Day, UK North, this month, Mark Sailes shared the lessons he’s learned behind building resilient systems, preparing teams for disruption and making sure hard-won experience benefits everyone.
Resilience planning tends to focus on the things we know can go wrong. We map dependencies, identify failure points and design systems to withstand disruption. But what happens when the failure that takes you offline is something nobody thought to plan for? In my experience, that's where the real lessons begin.
One of the defining moments in my own resilience journey happened in a former role.
Our data centre was based in Guernsey. A fishing boat dropped its anchor, and it dragged along the seabed severing all three undersea fibre optic cables connecting the island. Our entire data centre was suddenly cut off.
This is a good example of the kind of event that makes resilience planning so challenging. You can plan for component failures, software bugs and network outages, but it's much harder to anticipate every combination of circumstances that could take a critical service offline.
At the time, our disaster recovery arrangements relied on a cold standby on the mainland. It was there for exactly this kind of situation but having a backup environment and knowing that it will work when you need it are two very different things.
After the incident, we spent the following year carrying out fortnightly resilience tests to make sure the failover would actually work. That was an enormous amount of operational effort just to maintain confidence in our ability to recover.
Eventually, our team decided to take a different approach. We rebuilt on AWS using native regional services, embedding resilience into the architecture rather than relying on a cold standby and repeated manual fail overexercises. The result was a more resilient solution, less operational toil and, importantly, more time to improve the system rather than continually rehearse its recovery.
This experience shaped how I think about resilience. You can't anticipate every failure, but you can change how you design systems, how you prepare people and how you learn when something goes wrong.
When an incident happens, there's often one engineer who gets to the bottom of it. They spend hours investigating, find the root cause, implement a fix and get the service back up. Everyone is relieved, the incident is closed and the team moves on.
But what happens to everything that engineer learned at three o'clock in the morning, especially if they ultimately leave the business? Too often, the knowledge stays with the person who fixed the problem. The next engineer encounters the same failure months later and has to work it all out again. We need to make sure that doesn't happen.
One way to do this is through game days, these are structured exercises in which teams deliberately recreate failure scenarios and practice responding to them in a safe environment. After resolving an incident, ask the engineer who fixed it to lead a game day. Give the wider team their own temporary environment in which they can reproduce the failure, investigate the symptoms and understand why the original problem occurred.
The objective isn't simply to teach everyone how to apply the same fix. It's to help them understand the failure mode and learn how to avoid introducing it in the first place.
By practising in an environment where nothing real is at risk, people can ask questions, experiment and build confidence without the pressure of a live customer incident. It's a practical way to turn individual experience into collective capability.
There's another problem with testing resilience; your own team knows its system better than anyone else. That's a strength, but it can also create blind spots.
When you know how something is supposed to work, it's easy to make assumptions about how it might fail. You naturally test the scenarios you've already considered.
So, try inviting people from outside your team to challenge those assumptions.
In my previous roles, I would bring in technical leads and site reliability engineers (SREs) from other teams and ask them to find ways to break the system. Ideally, they'd have enough freedom to explore failure scenarios the team hasn't anticipated, including tests that are not announced in advance.
This is a form of red teaming, deliberately challenging a system to discover weaknesses before they affect customers.
People who haven't built or maintained your system bring a different perspective. They may question decisions your team takes for granted or identify dependencies nobody has considered.
It also gives your on-call engineers the opportunity to practice diagnosing problems and responding to unexpected behaviour. The key is to make these exercises safe, controlled and focused on learning. The aim isn't to catch someone out, it's to uncover weaknesses while there's still time to address them.
Most organisations are good at celebrating commercial success. We share the new customer, the major deal and the successful launch.
But what about the engineer who eliminates a recurring failure? Or the team that reduces recovery time, removes a fragile dependency or prevents a whole category of incidents? Those achievements deserve to be shared, too.
At AWS, the idea of an "ops win" email is a way to give significant operational improvements the same visibility as other business successes.
The important thing is to do it properly. Explain what the problem was, what changed and what measurable improvement followed. Make the information clear enough that another team, even one that wasn't involved, can understand the lesson and apply it.
Sharing operational improvements helps prevent teams from repeatedly solving the same problems in isolation. It also reinforces a culture in which improving reliability and reducing operational toil are recognised as valuable work.
Resilience isn't just about responding well when things break. It's also about recognising and rewarding the work that makes those failures less likely to happen again.
Of course, how we build systems still matters enormously. My experience with the Guernsey data centre reinforced the value of designing resilience into the architecture, rather than treating recovery as something to bolt on afterwards.
Multi-region architectures are one example. They can help reduce dependence on a single region, but they also introduce challenges that need careful consideration.
For systems that accept writes in multiple regions, for example, ensuring that records remain consistent and duplicate writes are handled correctly can add significant complexity.
Developments in AWS services are making some of these challenges easier to address. Amazon DynamoDB Global Tables with multi-region strong consistency (MRSC), alongside improvements in Amazon Time Sync Service, offer capabilities that can support multi-region designs.
The important point is that we should take advantage of the resilience capabilities available in the platforms we use, while understanding the trade-offs and operational implications of our chosen architecture.
Technology can remove entire categories of manual work, but it doesn't remove the need for teams to understand how their systems behave when something goes wrong. The strongest approach combines resilient architecture with people who know how to operate it.
If there's one message I'd take from these experiences, it's that resilience needs to be an ongoing practice, not a document that gets reviewed once a year.
Here are five practical things teams can start doing:
1. Build resilience in from the start. Consider failure scenarios, dependencies and recovery requirements when designing a service, rather than relying on recovery plans to compensate for fragile architecture.
2. Turn incidents into game days. Recreate significant failures in safe, temporary environments so the whole team learns from the experience, not just the person who resolved the incident.
3. Invite outsiders to challenge your assumptions. Ask people from other teams to test your systems and explore failure scenarios you might never have considered.
4. Share and celebrate operational improvements. Document measurable wins and circulate the lessons so other teams can benefit without having to discover the same problems themselves.
5. Use the resilience capabilities your platform provides. Look for opportunities to automate recovery, reduce manual operational effort and simplify complex resilience arrangements.
You don't need to transform everything at once. Start with one critical workload, identify where the biggest risks lie and choose one practical improvement to make. Then test it, learn from it and build from there.
Ultimately, resilience isn't just about having a backup or ticking a box in a disaster recovery plan. It's about creating the conditions in which people can respond effectively when the unexpected happens and ensuring that every incident leaves us better prepared for the next one.
If you'd like to explore how to build resilience into your projects from the outset, or how to prepare your teams for the failures they haven't yet imagined, we’d love to hear from you. Contact the team here.

What happens when the failure you never imagined actually happens? At the inaugural AWS Resiliency Day, UK North, this month, Mark Sailes shared the lessons he’s learned behind building resilient systems, preparing teams for disruption and making sure hard-won experience benefits everyone.
Resilience planning tends to focus on the things we know can go wrong. We map dependencies, identify failure points and design systems to withstand disruption. But what happens when the failure that takes you offline is something nobody thought to plan for? In my experience, that's where the real lessons begin.
One of the defining moments in my own resilience journey happened in a former role.
Our data centre was based in Guernsey. A fishing boat dropped its anchor, and it dragged along the seabed severing all three undersea fibre optic cables connecting the island. Our entire data centre was suddenly cut off.
This is a good example of the kind of event that makes resilience planning so challenging. You can plan for component failures, software bugs and network outages, but it's much harder to anticipate every combination of circumstances that could take a critical service offline.
At the time, our disaster recovery arrangements relied on a cold standby on the mainland. It was there for exactly this kind of situation but having a backup environment and knowing that it will work when you need it are two very different things.
After the incident, we spent the following year carrying out fortnightly resilience tests to make sure the failover would actually work. That was an enormous amount of operational effort just to maintain confidence in our ability to recover.
Eventually, our team decided to take a different approach. We rebuilt on AWS using native regional services, embedding resilience into the architecture rather than relying on a cold standby and repeated manual fail overexercises. The result was a more resilient solution, less operational toil and, importantly, more time to improve the system rather than continually rehearse its recovery.
This experience shaped how I think about resilience. You can't anticipate every failure, but you can change how you design systems, how you prepare people and how you learn when something goes wrong.
When an incident happens, there's often one engineer who gets to the bottom of it. They spend hours investigating, find the root cause, implement a fix and get the service back up. Everyone is relieved, the incident is closed and the team moves on.
But what happens to everything that engineer learned at three o'clock in the morning, especially if they ultimately leave the business? Too often, the knowledge stays with the person who fixed the problem. The next engineer encounters the same failure months later and has to work it all out again. We need to make sure that doesn't happen.
One way to do this is through game days, these are structured exercises in which teams deliberately recreate failure scenarios and practice responding to them in a safe environment. After resolving an incident, ask the engineer who fixed it to lead a game day. Give the wider team their own temporary environment in which they can reproduce the failure, investigate the symptoms and understand why the original problem occurred.
The objective isn't simply to teach everyone how to apply the same fix. It's to help them understand the failure mode and learn how to avoid introducing it in the first place.
By practising in an environment where nothing real is at risk, people can ask questions, experiment and build confidence without the pressure of a live customer incident. It's a practical way to turn individual experience into collective capability.
There's another problem with testing resilience; your own team knows its system better than anyone else. That's a strength, but it can also create blind spots.
When you know how something is supposed to work, it's easy to make assumptions about how it might fail. You naturally test the scenarios you've already considered.
So, try inviting people from outside your team to challenge those assumptions.
In my previous roles, I would bring in technical leads and site reliability engineers (SREs) from other teams and ask them to find ways to break the system. Ideally, they'd have enough freedom to explore failure scenarios the team hasn't anticipated, including tests that are not announced in advance.
This is a form of red teaming, deliberately challenging a system to discover weaknesses before they affect customers.
People who haven't built or maintained your system bring a different perspective. They may question decisions your team takes for granted or identify dependencies nobody has considered.
It also gives your on-call engineers the opportunity to practice diagnosing problems and responding to unexpected behaviour. The key is to make these exercises safe, controlled and focused on learning. The aim isn't to catch someone out, it's to uncover weaknesses while there's still time to address them.
Most organisations are good at celebrating commercial success. We share the new customer, the major deal and the successful launch.
But what about the engineer who eliminates a recurring failure? Or the team that reduces recovery time, removes a fragile dependency or prevents a whole category of incidents? Those achievements deserve to be shared, too.
At AWS, the idea of an "ops win" email is a way to give significant operational improvements the same visibility as other business successes.
The important thing is to do it properly. Explain what the problem was, what changed and what measurable improvement followed. Make the information clear enough that another team, even one that wasn't involved, can understand the lesson and apply it.
Sharing operational improvements helps prevent teams from repeatedly solving the same problems in isolation. It also reinforces a culture in which improving reliability and reducing operational toil are recognised as valuable work.
Resilience isn't just about responding well when things break. It's also about recognising and rewarding the work that makes those failures less likely to happen again.
Of course, how we build systems still matters enormously. My experience with the Guernsey data centre reinforced the value of designing resilience into the architecture, rather than treating recovery as something to bolt on afterwards.
Multi-region architectures are one example. They can help reduce dependence on a single region, but they also introduce challenges that need careful consideration.
For systems that accept writes in multiple regions, for example, ensuring that records remain consistent and duplicate writes are handled correctly can add significant complexity.
Developments in AWS services are making some of these challenges easier to address. Amazon DynamoDB Global Tables with multi-region strong consistency (MRSC), alongside improvements in Amazon Time Sync Service, offer capabilities that can support multi-region designs.
The important point is that we should take advantage of the resilience capabilities available in the platforms we use, while understanding the trade-offs and operational implications of our chosen architecture.
Technology can remove entire categories of manual work, but it doesn't remove the need for teams to understand how their systems behave when something goes wrong. The strongest approach combines resilient architecture with people who know how to operate it.
If there's one message I'd take from these experiences, it's that resilience needs to be an ongoing practice, not a document that gets reviewed once a year.
Here are five practical things teams can start doing:
1. Build resilience in from the start. Consider failure scenarios, dependencies and recovery requirements when designing a service, rather than relying on recovery plans to compensate for fragile architecture.
2. Turn incidents into game days. Recreate significant failures in safe, temporary environments so the whole team learns from the experience, not just the person who resolved the incident.
3. Invite outsiders to challenge your assumptions. Ask people from other teams to test your systems and explore failure scenarios you might never have considered.
4. Share and celebrate operational improvements. Document measurable wins and circulate the lessons so other teams can benefit without having to discover the same problems themselves.
5. Use the resilience capabilities your platform provides. Look for opportunities to automate recovery, reduce manual operational effort and simplify complex resilience arrangements.
You don't need to transform everything at once. Start with one critical workload, identify where the biggest risks lie and choose one practical improvement to make. Then test it, learn from it and build from there.
Ultimately, resilience isn't just about having a backup or ticking a box in a disaster recovery plan. It's about creating the conditions in which people can respond effectively when the unexpected happens and ensuring that every incident leaves us better prepared for the next one.
If you'd like to explore how to build resilience into your projects from the outset, or how to prepare your teams for the failures they haven't yet imagined, we’d love to hear from you. Contact the team here.