Alerts in the Experimental alerting system
In the experimental alerting system, Kibana tracks each problem as an alert episode, the full lifecycle of a condition from first detection through recovery. Kibana writes a rule event for each matching row, and the episode is the grouping of those events that share an episode.id.
This page explains the core concepts you need to work with the experimental alerting system: how alert episodes move through lifecycle states, and how series group episodes over time for the same monitored subject.
Every alert episode moves through these states:
inactive → pending → active → recovering → inactive
| State | What it means |
|---|---|
| Inactive | Problem fully resolved. An action policy can invoke a workflow to send a recovery notification. |
| Pending | Errors detected, but the system is waiting to confirm it's a real problem before the episode becomes active. |
| Active | Problem confirmed and ongoing. An action policy can evaluate the episode and invoke a workflow. |
| Recovering | Errors have stopped, but the system is waiting to confirm it's truly resolved. |
Example: A checkout-latency episode moving through all four states
A checkout-latency rule runs every 5 minutes. It has an activation threshold of 2 consecutive breaches and a recovery threshold of 2 consecutive clears. The episode opens only after consecutive breaches meet the activation threshold and closes only after consecutive clears meet the recovery threshold. The system waits for confirmation in both directions. An action policy matches this rule's episodes and invokes a workflow that notifies on-call when the episode becomes active and when it recovers.
- 14:00: Routine check. p95 is within budget. No episode exists yet. The series is
inactive. - 14:05: p95 jumps to 3.1s. The rule detects the first breach. Kibana writes a rule event, opens the episode in
pending, and starts counting consecutive breaches. - 14:10: p95 is still elevated. The second consecutive breach meets the activation threshold. The episode moves from
pendingtoactive. The action policy invokes the workflow, which pages the engineer. - 14:10–14:45: Every evaluation finds high latency. The episode stays
active. Kibana doesn't open new episodes. One episode tracks one problem, no matter how many times the rule evaluates while the condition holds. - 14:50: p95 drops back under 2s. The first clean check moves the episode from
activetorecovering. The system starts counting consecutive clears. - 14:55: A second consecutive clear meets the recovery threshold. The episode moves from
recoveringtoinactive. The engineer receives a recovery notification.
What this illustrates:
inactiveis the resting state. The series exists but isn't tracking a problem.pendingis the confirmation gate on the way in. Without it, a brief latency spike at 14:05 opens and immediately closes an episode, creating noise. The threshold filters that out.activeis the steady state of an ongoing problem. The episode accumulates evaluations without branching, covering the entire outage from first confirmation to first clear.recoveringis the confirmation gate on the way out. Without it, a single good evaluation at 14:50 closes the episode, even if latency bounces back up at 14:55. The threshold prevents premature resolution.inactiveagain signals confirmed recovery. The episode closes and the recovery notification fires only after the condition has cleared consistently.
A series is the ongoing relationship between a rule and one specific thing it monitors. It exists for as long as that rule keeps monitoring that thing, and can contain many alert episodes over its lifetime, one for each time that thing had a problem.
Think of it like a patient's medical file. The file persists as long as the patient is in the system. Individual health incidents come and go, but the file stays. Each incident is an episode in the same series.
Snooze operates at the series level, not the alert episode level. If you snooze checkout-service, you're silencing all notifications from that series for the next X hours, regardless of how many new alert episodes start during that time.
From here, you can view, manage, and query alert episode data, and query .rule-events in Discover.
- View and manage alerts: Open the alert episodes table, triage active episodes, and acknowledge, snooze, or resolve them.
- Rule events: What Kibana writes to
.rule-eventsand how those events form an episode. - Rule event data model: Where rule events are stored and how they differ by
type. - Query experimental alerting system alert history in Discover: Use ES|QL to query
.rule-eventsand.alert-actionsfor exploratory analysis and dashboards. - Query signals: Query events with
type: signalin Discover and use them as input to a rule that opens an episode.
Because the experimental alerting system is still evolving, its UI can change before general availability. Rather than pointing to an exact button or menu, the documentation focuses on the underlying concepts and behavior. If something doesn't match what you see in the Kibana UI, look for the closest equivalent instead. The concepts and behaviors described in the documentation still apply.