IncidentFlare
← All guides

Alerts and escalation in 5 minutes

A status page that nobody is watching is a diary, not an alert. This is how to make a failing health check reach a human.

1. Add a channel

In the dashboard open AlertsAdd channel. Four types work today:

Hit Send test straight away. If the channel is misconfigured you’ll see the exact error instead of finding out during an outage.

2. Or page a person, not a place

Under On-call you can add the people who can be alerted and put them in a rotation that hands over daily or weekly. An escalation step can then point at the rotation instead of a channel, and we page whoever is holding it through every contact point they have. Somebody away for a weekend gets a cover window, and the rotation picks up again when it ends.

The rotation stores a timezone name, not an offset, so a 09:00 handover stays at 09:00 through daylight saving instead of drifting an hour twice a year.

3. Order the steps

Escalation steps run from the moment the alert opens. A sane starting point for a two-person team:

  1. Step 1 — Slack, delay 0 minutes.
  2. Step 2 — your email, delay 10 minutes.
  3. Step 3 — co-founder’s email or a webhook to your phone, delay 25 minutes.

Acknowledging the alert stops every remaining step, so the 3am escalation only happens when nobody picked it up. Every page carries an acknowledge link, so stopping the escalation is one tap from the Slack message or the email — no dashboard, no login.

4. Decide what happens when the steps run out

Two settings sit under the steps, and both default to on:

Resolving the alert is what ends it for good. The timeline on the Alerts page records every repeat and every revival, so you can see what actually happened after the fact.

5. Run a drill

Use Fire test alert. It creates a real alert, walks the real escalation policy, and records each delivery — the same path a real outage takes. Acknowledge it when it reaches you, and you know paging works.

What triggers alerts

Two consecutive failed checks on a monitor. That threshold avoids paging on a single blip. The same event also opens a public incident and flips the linked component, and a recovery resolves both the incident and the alert.

Start freeGo to dashboard — paging is included on the free plan.