Skip to main content

Incidents and investigations

With Incident grouping on, a pattern tracks each problem as one incident from its first bad run to its recovery, instead of raising a separate alert every run. This page follows one incident run by run; Alerts and subscriptions then shows the same moments from the subscriber's side.

Three words​

WordMeansIn the example below
SegmentOne combination of dimension values. Some teams say series. A pattern with no dimensions has one segment: the whole KPI.Store: NYC
EpisodeOne segment's run of out-of-band runs, from the first bad run until it closes. It opens when enough recent runs are out of band and the Impact rules hold, and closes after the configured number of normal runs or when the Impact rules stop holding. A gap inside the Trigger window does not break it. A segment has at most one open episode at a time.NYC, out of band from Jul 6 09:00, closed 17:00
IncidentThe episodes the grouping mode puts together, live from its first episode's open until its last live episode closes. One incident is one Triage page, one Slack or Teams message, one email conversation.NYC and SF together: opened 11:00, resolved 19:00

What every run does​

StepWhat happens
DetectEach segment's value is compared with its baseline.
Update episodesA segment opens an episode once Out-of-band runs to open of the runs in its Trigger window are out of band and the Impact rules hold over the stretch so far. A segment still out of band advances its episode. In-band runs to close normal runs in a row close it, and so do Impact rules that stop holding.
GroupAn opened episode joins an open incident by the grouping mode, or opens a new one. An incident resolves when its last live episode closes.
AlertSubscriptions update or send by the rules on the alerts page.
InvestigateWith How should Bicycle investigate this pattern? set to Automatically, the analysis runs for every incident that moved this run.

Out of band means outside the band, whether or not that run's own alert was suppressed by Fire alert when. A segment whose episode did not move changes nothing: no update, no alert, no analysis.

An incident's life​

MomentThe incidentThe investigationAlerts
Opens: a segment's episode qualifies and no open incident takes itCreated, ongoing, one segmentRunsMessage posted; email by selection
Advances: the segment is still out of bandNumbers and window updatedRe-runs over the grown windowMessage updated; no email
A segment joinsOne more live episodeRe-runs for the whole incidentMessage updated; email without a summary
A segment recoversIts episode closes; the incident stays ongoingRe-runsMessage updated; email without a summary
Resolves: the last live episode closesResolved, with how long it lastedRe-runs once more over the closed windowsMessage shows Resolved; email without a summary
Season cut: an episode reaches one season of the baselineThe episode is cut and continues as a new one. In the default grouping mode the continuation opens a new incident.Runs for the new incidentThe cut sends nothing; the continuation alerts as an opening
Nothing movedUnchangedNothingNothing

"Email without a summary" is the Custom selection with no summary block. Selections with a summary email once, when the first analysis completes. The alerts page has the full table.

When the investigation runs​

  • Only with Automatically. On demand runs nothing on its own; you start it with Run RCA.
  • The trigger is movement on that run, not "the incident is still open". A segment still out of band, a segment joining, and a segment recovering each re-run the analysis. A run where nothing moved runs nothing.
  • Every re-run covers the whole incident, each segment over its own episode window, and replaces the previous analysis in Triage and on the Slack or Teams message. Email stays quiet after its first analysis email.
  • One analysis at a time per incident. A run that arrives while one is going is skipped; the next run that moves the incident starts a fresh one.
  • With an RCA start date set, only runs whose period ends on or after that date are analysed.
  • Run RCA analyses the incident as it stands and sends no alerts.
  • On an hourly pattern, a problem that lasts all day is analysed every hour. Without incident grouping, each run is analysed once, on its own.

Worked example​

The same incident as on the alerts page: an hourly pattern, Drop in Orders, by Store, default grouping mode, investigating Automatically.

TimeSegmentsEpisodesIncidentInvestigation
09:00NYC dropsOut of band; the open rule is not met yet
11:00NYC still downNYC episode opens, dated from 09:00Opens, one segmentRuns, completes 11:05
12:00 to 14:00NYC still downAdvancesNumbers updateRe-runs each hour
15:00SF dropsSF episode opensSF joins, two segmentsRe-runs for both segments
17:00NYC recoversNYC episode closesOngoing, one of two liveRe-runs
19:00SF recoversSF episode closesResolves, lasted 8hFinal run over the closed windows
Next daySF drops againNew SF episodeNew incidentRuns

The alerts page shows two of these analyses, at 11:05 and 15:06, to keep its tables short; every other moved run re-runs it too.

Next: Alerts and subscriptions — routing what an incident produces to Slack, Teams, or email.