Incidents and investigations
With Incident grouping on, a pattern tracks each problem as one incident from its first bad run to its recovery, instead of raising a separate alert every run. This page follows one incident run by run; Alerts and subscriptions then shows the same moments from the subscriber's side.
Three words
| Word | Means | In the example below |
|---|---|---|
| Segment | One combination of dimension values. Some teams say series. A pattern with no dimensions has one segment: the whole KPI. | Store: NYC |
| Episode | One segment's run of out-of-band runs, from the first bad run until it closes. It opens when enough recent runs are out of band and the Impact rules hold, and closes after the configured number of normal runs or when the Impact rules stop holding. A gap inside the Trigger window does not break it. A segment has at most one open episode at a time. | NYC, out of band from Jul 6 09:00, closed 17:00 |
| Incident | The episodes the grouping mode puts together, live from its first episode's open until its last live episode closes. One incident is one Triage page, one Slack or Teams message, one email conversation. | NYC and SF together: opened 11:00, resolved 19:00 |
What every run does
| Step | What happens |
|---|---|
| Detect | Each segment's value is compared with its baseline. |
| Update episodes | A segment opens an episode once Out-of-band runs to open of the runs in its Trigger window are out of band and the Impact rules hold over the stretch so far. A segment still out of band advances its episode. In-band runs to close normal runs in a row close it, and so do Impact rules that stop holding. |
| Group | An opened episode joins an open incident by the grouping mode, or opens a new one. An incident resolves when its last live episode closes. |
| Alert | Subscriptions update or send by the rules on the alerts page. |
| Investigate | With How should Bicycle investigate this pattern? set to Automatically, the analysis runs for every incident that moved this run. |
Out of band means outside the band, whether or not that run's own alert was suppressed by Fire alert when. A segment whose episode did not move changes nothing: no update, no alert, no analysis.
An incident's life
| Moment | The incident | The investigation | Alerts |
|---|---|---|---|
| Opens: a segment's episode qualifies and no open incident takes it | Created, ongoing, one segment | Runs | Message posted; email by selection |
| Advances: the segment is still out of band | Numbers and window updated | Re-runs over the grown window | Message updated; no email |
| A segment joins | One more live episode | Re-runs for the whole incident | Message updated; email without a summary |
| A segment recovers | Its episode closes; the incident stays ongoing | Re-runs | Message updated; email without a summary |
| Resolves: the last live episode closes | Resolved, with how long it lasted | Re-runs once more over the closed windows | Message shows Resolved; email without a summary |
| Season cut: an episode reaches one season of the baseline | The episode is cut and continues as a new one. In the default grouping mode the continuation opens a new incident. | Runs for the new incident | The cut sends nothing; the continuation alerts as an opening |
| Nothing moved | Unchanged | Nothing | Nothing |
"Email without a summary" is the Custom selection with no summary block. Selections with a summary email once, when the first analysis completes. The alerts page has the full table.
When the investigation runs
- Only with Automatically. On demand runs nothing on its own; you start it with Run RCA.
- The trigger is movement on that run, not "the incident is still open". A segment still out of band, a segment joining, and a segment recovering each re-run the analysis. A run where nothing moved runs nothing.
- Every re-run covers the whole incident, each segment over its own episode window, and replaces the previous analysis in Triage and on the Slack or Teams message. Email stays quiet after its first analysis email.
- One analysis at a time per incident. A run that arrives while one is going is skipped; the next run that moves the incident starts a fresh one.
- With an RCA start date set, only runs whose period ends on or after that date are analysed.
- Run RCA analyses the incident as it stands and sends no alerts.
- On an hourly pattern, a problem that lasts all day is analysed every hour. Without incident grouping, each run is analysed once, on its own.
Worked example
The same incident as on the alerts page: an hourly pattern, Drop in Orders, by Store, default grouping mode, investigating Automatically.
| Time | Segments | Episodes | Incident | Investigation |
|---|---|---|---|---|
| 09:00 | NYC drops | Out of band; the open rule is not met yet | ||
| 11:00 | NYC still down | NYC episode opens, dated from 09:00 | Opens, one segment | Runs, completes 11:05 |
| 12:00 to 14:00 | NYC still down | Advances | Numbers update | Re-runs each hour |
| 15:00 | SF drops | SF episode opens | SF joins, two segments | Re-runs for both segments |
| 17:00 | NYC recovers | NYC episode closes | Ongoing, one of two live | Re-runs |
| 19:00 | SF recovers | SF episode closes | Resolves, lasted 8h | Final run over the closed windows |
| Next day | SF drops again | New SF episode | New incident | Runs |
The alerts page shows two of these analyses, at 11:05 and 15:06, to keep its tables short; every other moved run re-runs it too.
Next: Alerts and subscriptions — routing what an incident produces to Slack, Teams, or email.