GMax Mart
Home Services Pricing Portfolio FAQ Reviews Blog Support Careers Change Language

Website Maintenance

Your website is down: a step-by-step incident response plan

· 7 min read

Your website is down: a step-by-step incident response plan

The first person to notice is almost never you. A customer calls, a salesperson forwards a screenshot, and suddenly three people are logged into the hosting panel clicking different buttons. Most outages get longer, not shorter, in those first ten minutes. Having a written plan for when your website is down is what turns a scramble into a sequence.

What follows is a runbook you can adapt: confirm, diagnose, escalate, communicate, recover, review. Print it, put it in your team's shared drive, and make sure the phone numbers in it are current.

Step one: confirm the outage is real

Before alerting anyone, spend two minutes establishing what is actually broken. A surprising share of reported outages turn out to be one person's browser cache, office Wi-Fi or a local DNS problem.

  • Open the site on mobile data with Wi-Fi switched off.
  • Try it in a private browsing window, and on a second device.
  • Use an external checker or an uptime service to test from outside your network.
  • Test both the bare domain and the www version, and both http and https.
  • Check whether one page is broken or the whole site, since a failing checkout is an incident too even when the home page loads.

Note the exact time you confirmed it. Every later step, including any hosting claim, depends on having a timeline.

Step two: read the error, because it names the culprit

The message on screen narrows the cause faster than any guess. Learn to recognise the common ones.

Nothing loads at all

A connection timeout or a DNS error usually points away from your application. A DNS failure suggests nameserver or registrar trouble, possibly an expired domain. A refused connection suggests the server or the web server process is down.

A browser security warning

A certificate warning means the site is running but the certificate has expired, does not match the hostname, or is serving an incomplete chain. Visitors see a red screen and leave, so treat it as a full outage.

A 500 or 503 page

The server answered, so DNS and networking are fine. A 500 is typically application code, a failed deployment or a database connection problem. A 503 often means resource exhaustion or a service that has stopped.

A hosting suspension notice

Usually billing, resource abuse or a malware detection. Check your registered email, including spam, before assuming a technical fault.

The site loads but is blank or half-rendered

Often a PHP fatal error with display turned off, a memory limit hit, or a broken asset path after a change. The server error log gives the answer in one line.

Step three: escalate to the right person with the right information

Decide in advance who owns which layer, and keep the list somewhere you can reach from a phone: hosting support, your developer or agency, the domain registrar, and the payment gateway if transactions are affected.

When you raise a ticket, include everything at once instead of trading messages for an hour:

  1. Domain name and server IP
  2. The exact error text or a screenshot with the URL visible
  3. When it started, in a stated time zone, and whether it is constant or intermittent
  4. What changed recently: a deployment, a plugin update, a DNS edit, a traffic spike
  5. What you have already tested, so nobody repeats it
  6. Output of a basic check such as a ping or traceroute if you can run one

State the business impact plainly, for example that checkout is unavailable and ads are running. Support queues are triaged, and a clear impact statement is not exaggeration.

Step four: tell customers before they have to ask

Silence during an outage reads as incompetence. A short, honest message buys you a great deal of patience.

Prepare three things in advance so they are ready to send:

  • A holding message for social media or WhatsApp broadcast: what is affected, that you are working on it, and how else to reach you.
  • A one-line script for whoever answers the phone, including an alternative way to place an order or log an enquiry.
  • A maintenance page on a separate host or platform, so customers reach something branded rather than a blank error.

Pause paid campaigns as soon as you confirm the outage. Ads keep spending while the destination is broken, and restarting them later takes seconds. If you take orders on WhatsApp or by phone, say so loudly; a proportion of the demand can be captured rather than lost.

Do not promise a restoration time you cannot control. "We expect an update within an hour" is better than "back in ten minutes".

Step five: recover without making it worse

Incidents get extended by well-meant improvisation. A few rules keep the damage contained.

Change one thing at a time and test after each change. Two simultaneous edits mean you will never know which one helped.

Before restoring a backup over a live database, take a copy of the current state first, even if it looks broken. Orders placed since the backup exist only in that copy.

Keep a running note with timestamps of every action taken and by whom. It takes seconds during the incident and saves hours afterwards.

If the trigger was a recent deployment, rolling back is almost always faster than fixing forward under pressure. Get the site up on the previous version, then debug calmly.

Resist the urge to delete things. Disabling a suspect plugin or module is reversible; deleting files is not.

Step six: the review that stops a repeat

Within a day or two, while details are fresh, write half a page covering four points: what happened, the timeline from first symptom to resolution, the actual root cause rather than the immediate trigger, and what will change.

Keep the actions concrete and assign each one an owner and a date. Typical outcomes are useful and unglamorous: add external uptime monitoring with alerts to two people, set certificate and domain expiry reminders, move backups off the same server, raise a resource limit, or add a staging environment so deployments are tested before they reach customers.

Track how long detection took separately from how long the fix took. Most small businesses discover that detection, not repair, is the bulk of their downtime, and that is the cheapest part to improve.

Write your runbook before you need it

Put the contact list, the login locations, the escalation order and the holding message into one document today, and give it to everyone who might be first to hear about an outage. The value is not in the detail; it is in nobody having to improvise at 9pm.

If your outages keep tracing back to the server rather than the code, the underlying platform may be the problem. Our team can look at your logs through the support desk and advise whether a move to managed Linux hosting would remove the recurring cause.

Frequently asked questions

Who should I call first, my host or my developer?

Let the symptom decide. DNS failures, refused connections and suspensions go to the host. Application errors, blank pages and problems right after a change go to your developer. If you genuinely cannot tell, contact both with the same information.

How quickly should I be told about an outage?

Aim for minutes, not hours. An external monitor checking every one to five minutes, with alerts by email and a messaging channel to at least two people, covers the case where the first person is unreachable.

Is it worth putting up a maintenance page during an incident?

Yes, if you can serve one, because it looks deliberate and lets you give an alternative contact route. Host it somewhere independent of the failing server, otherwise it will be unavailable exactly when you need it.

Should I post about an outage on social media?

If customers are affected and asking, a short factual update is better than silence. Keep it brief, avoid technical blame, and post again when service is restored so the last message people see is the resolution.

Thinking about a website?

See what a package covers and what it costs, or ask us about your own project.

Read next

Thinking…