ARTICLE 04 · RELIABILITY · 2026-09-06

When the system goes down at peak sales: the monitoring, backups and recovery plan to have ready before that day

The damage from an outage is decided long before the day it happens, by three questions. Has anyone measured how many users the system can handle? Will anyone spot a problem before customers do? And has anyone ever actually tried restoring the backups you have?

When the system goes down at peak sales: the monitoring, backups and recovery plan to have ready before that day
The short version
  • Outages come in three kinds: the system can't handle the traffic, data gets lost, or something breaks quietly without anyone knowing. Each one needs different preparation.
  • A backup you have never tried restoring doesn't count as a backup yet.
  • Business owners have to answer two questions themselves: how long can the system be down, and how many hours of data can you afford to lose? The technical team designs around those answers.

Three kinds of outage businesses run into

It is 8 p.m. sharp on campaign day, and thousands of customers hit the website at once. Pages start spinning, the marketing team asks in the group chat what is going on, and nobody can answer. This is the kind of outage people picture first, but the kind that does more damage is usually quieter.

The first kind is being unable to handle the load. A system that copes easily with normal traffic slows down until it is unusable when more people arrive at once than it has ever seen. The bottleneck is usually the database, and adding more web servers won't help.

The second kind is data loss. The causes range from a staff member deleting the wrong table and hardware failure to ransomware that encrypts every file. That last one will spread to the backups as well if they are permanently connected to the main system.

The third kind is a silent failure. The website still opens, but some key steps have stopped working. Customers pay successfully but the order isn't recorded, or the confirmation email never goes out. This is the most expensive kind, because hours pass before anyone notices, and every transaction then has to be fixed one at a time.

Three kinds of outage: too much traffic, lost data, and silent failures nobody notices
Three kinds of outage: too much traffic, lost data, and silent failures nobody notices

Apps built with vibe code come with every screen in place, but no monitoring, no data backups and nobody who gets woken up at 2 a.m. Nobody told the AI to build those things, and during the demo nobody saw they were missing.

Four layers of readiness, and the evidence to ask for

A restaurant that is ready for Friday night knows how many dishes the kitchen can turn out per hour, has someone watching the front of house, has a backup stove for when the gas runs out, and everyone knows who makes the call when the power goes. An online system needs the same four things, and each one comes with evidence a business owner can ask to see without understanding the technology.

Readiness layerThe business owner's questionEvidence to ask for
CapacityHow many users can the system handle at the same time before it slows down?The latest load test results, with the figures and the test date
VisibilityIf something went wrong right now, who would know first, our customers or us?The monitoring dashboard and the list of alerts that have been set up
RecoveryIf the database were lost now, up to what point could we get the data back, and how many hours would it take?The record of the most recent data recovery drill
ResponseIf the system goes down on a Saturday night, who decides, and who tells customers?A runbook and an on-call roster with real names on it

The recovery layer needs two answers from the business owner. The first is how long the system can be down before the damage to the business becomes unacceptable. Engineers call this the RTO. The second is how far back you can afford to lose data, called the RPO. A shop taking dozens of orders a minute can't accept an RPO of one day, because that would mean losing a whole day of orders, while a company website updated once a month can live with it easily. The shorter you set these two values, the higher the cost, so the numbers should come from the business.

The long-standing rule for backups is 3-2-1: keep three copies, on two types of media, with one copy off-site. The off-site copy that isn't connected to the main system is the one that survives a ransomware attack.

For the development team · Technical detail

Set SLOs for the order and payment path. Alert on the symptoms customers experience, such as error rate, P95 latency and payment success rate. Run synthetic checks that simulate an order every few minutes. Use point-in-time backups with automated restore tests, load test at levels above the expected peak, and take orders through a queue so that transactions aren't lost when the database slows down.

A cosmetics brand and the first twenty-five minutes of its campaign

A cosmetics brand that sells through its own website launches a campaign at 8 p.m. At 8:04 the site starts to slow. At 8:10 the checkout page freezes. Some customers are charged but never get an order number, because the bank's confirmation signal arrives while the system isn't responding. The team finds out from comments on the brand's page at 8:25 and manages to restart the server at 8:40. The next morning, three staff have to match the bank's charges against the orders in the system one by one.

The team learns about the problem from customer comments, twenty-five minutes after the campaign starts
The team learns about the problem from customer comments, twenty-five minutes after the campaign starts

Look back at what each layer of readiness would have changed. A load test a week before the campaign would have shown that the database was the bottleneck. Alerts tied to the payment success rate would have told the team within two minutes. An order queue would have held the bank's signals until the system came back. And a runbook would have said who posts the announcement on the page, so nobody had to ask in the middle of the incident.

Run a recovery drill every quarter

  1. Pick the latest backup of the main database.
  2. Restore it onto a separate machine, away from the live system, and time it from the start until it is usable.
  3. Check three things: what time the latest order in the backup is from, whether the number of customers matches the live system, and whether you can log in and open orders.
  4. Write down the time it took and the window of data that was lost, then compare them with the targets the business owner set.
  5. Fix whatever got stuck, then schedule the next drill.

The evidence is a drill record with the date, the time taken and the name of the person who ran it. Where drills often come unstuck is that only the database has been backed up, while product images, documents customers uploaded and system settings live somewhere else that has never been backed up.

What level of readiness is worth it for your business

Shops that sell through a marketplace or a ready-made store platform hardly need to think about this, because the provider already takes care of it. This article is about businesses with their own systems, which have to decide for themselves how many layers to prepare.

A workable way to think about it is to estimate the cost of downtime per hour. Take peak-hour sales, add the labor cost of the team that has to clean up, add the customers who won't come back, and compare the total with the cost of each readiness layer. Businesses with low downtime costs should at least have backups that have been through a recovery drill, plus basic alerts, because these two cost little and prevent damage that can't be undone.

A sensible order of investment is to start with backups and recovery drills, then monitor the paths that touch money, then run load tests before big campaigns, write a runbook and rehearse incident response, and finally, for businesses that can never stop, add a standby system that takes over automatically, known as failover.

Run a real, timed recovery drill to find out whether the backups work and how long a restore takes
Run a real, timed recovery drill to find out whether the backups work and how long a restore takes

Move up to the next layer when

  • Peak-hour sales are higher than the cost of preparing that layer
  • A campaign is expected to bring more visitors than the system has ever handled
  • Customers have spotted an incident before the team did
  • You have contracts with corporate customers that commit you to service times

The numbers to track after investing are the time from an incident to the team knowing about it, the time from knowing to the system coming back, the number of incidents customers spotted first, and the result of the latest recovery drill. If the time it takes the team to notice an incident doesn't fall after monitoring is installed, the alerts are still tied to machine metrics instead of what customers experience. Fix that before buying more servers.

DNA MAKER · RELIABILITY ENGINEERING

Get your systems ready before the next big campaign

Your executives decide how much downtime the business can accept, marketing knows the campaign calendar, and customer service knows what customers ask when the system has problems. DNA Maker gathers answers from all three and turns them into targets for the system: the periods when it must not go down, the paths to watch from order through payment to order confirmation, and the chain of decisions when an incident happens.

Our team reviews the architecture, runs load tests and stress tests to find the bottlenecks, sets up monitoring and alerts on the paths that touch money, sets backup schedules with recovery drills, and writes the disaster recovery plan and runbook together with your team. For systems that need high availability, we design the scaling and the standby servers. If you have a big campaign coming up in the next two or three months, bring the campaign calendar and the notes from the last time the system had problems, and the work will start from the riskiest point.

Software engineering glossary

Use these terms to agree with the technical team how much the system has to withstand. The first two are numbers the business side has to set.

TermWhat it isA simple exampleWhat executives should ask the development team
RTOThe longest time the system is allowed to be down, from the moment of the incident until it is back in useA shop decides the checkout page must be back within thirty minutesHave we ever actually hit our RTO in a drill?
RPOHow much data you can accept losing when you have to restore from a backupBackups run every fifteen minutes, so if a restore is needed, the lost orders go back no further than the last fifteen minutesIf we had to restore now, what time would the latest data we got back be from?
BackupA copy of data stored separately from the main system, used to restore when the real data is damaged or lostThe order database is copied to another location every night, with thirty days of history keptWhere are the backups stored, who can access them, and when did we last try a restore?
Load TestSimulating a large number of users on the system at once, to measure how much it can take and where it gets stuckSimulating five thousand customers buying at the same moment, a week before the campaignAt how many users does the system start to slow down, and which part is the bottleneck?
RunbookA step-by-step guide for handling incidents you have planned for, naming who acts and who decidesWhen the checkout page goes down, the guide says who checks what, who posts on the page and what message to useWhen was this guide last used in a drill, and have the people on the on-call roster read it?
Disaster RecoveryThe plans and systems for restoring service after a major incident, such as a data center outage or data being encryptedWhen the cloud provider goes down across a whole region, the team brings the system up from a copy in another regionWhich incidents does this plan cover, and who has to be involved to set it in motion?
FailoverSwitching automatically to a standby server or system when the main one stops workingThe main database goes down and the system switches to the standby within a minute, with nothing for customers to doHas the switchover been tested on the real system, and is any data lost during the switch?
Try this tomorrow: Ask your system administrator one question: when did we last actually try restoring a backup, and how long did it take? If the answer is never, book a drill within this month.