Prometheus No-Data Alerts: Practical Policy
Design Prometheus alerts for missing metrics without paging on every scale-down. Use absent_over_time, scrape health, and explicit no-data policy.
Ben Ennis
Published September 22, 2026
A quiet dashboard can mean a healthy service, a workload that scaled down, or a monitoring path that stopped reporting. Prometheus cannot infer which one you mean. An alert expression that returns no series simply produces no alert instance, so a failed exporter can become a silent gap unless you model presence explicitly.
This is a policy problem before it is a PromQL problem. Decide what “no data” means for each metric, then choose a signal that matches the failure you need to detect. A queue worker may legitimately disappear when its queue is empty. A control-plane scrape target should not disappear without an owner knowing. A service-level indicator may need both a data-presence alert and a burn-rate alert.
The examples below use Prometheus rules and Alertmanager. They distinguish an empty query result, a missing member of a multi-dimensional set, and a failed scrape. That distinction keeps scale-downs quiet while still giving an on-call engineer a clear path when telemetry disappears unexpectedly.
Define the meaning of an empty result
Start with a short contract for each important metric:
| Metric situation | Likely meaning | First signal | Human action |
|---|---|---|---|
| No queue items and exporter still responds | Valid zero | Exporter health stays normal | No page |
| Exporter target disappears | Collection failure or target removal | up or target inventory alert |
Inspect discovery and rollout |
| Required metric absent for a window | Instrumentation or query contract broke | absent_over_time() |
Check exporter, labels, and deploy |
| One region disappears while others remain | Partial outage or decommission | Per-label presence check | Compare target inventory with ownership |
| Query returns an evaluation error | Prometheus or data-source failure | Rule evaluation/error signal | Restore query path before trusting alerts |
Write the expected lifecycle beside the metric definition. “This series exists with value zero when idle” is a stronger contract than “the dashboard usually shows zero.” If a metric is created only after the first event, an empty result may be normal during startup; alerting on its absence would be noise.
A useful test is to ask what a responder should conclude from one missing series. If the answer is “nothing, because the workload was intentionally removed,” do not page from the metric alone. If the answer is “we no longer know whether customers are protected,” route the signal to the team that owns the telemetry.
Prometheus’s querying basics describe an instant vector as a set of time series with one sample at an evaluation timestamp. An empty instant vector is therefore a valid query result, not an automatic error. The rule must turn that absence into a value if you want a condition to match it.
Use absent_over_time() for required metrics
PromQL provides two presence functions. absent() returns a one-element vector with value 1 when its input has no elements at the evaluation time. absent_over_time() does the same when a range vector has contained no samples for the selected period.
For a required exporter metric, the range form is usually safer because it tolerates a brief scrape gap:
absent_over_time(
app_build_info{job="checkout-exporter"}[10m]
)
Turn that expression into an alert only when the metric is required to exist for this job:
groups:
- name: telemetry.contracts
interval: 30s
rules:
- alert: CheckoutExporterMetricMissing
expr: |
absent_over_time(
app_build_info{job="checkout-exporter"}[10m]
)
for: 2m
labels:
severity: ticket
team: checkout
annotations:
summary: "Checkout exporter metric is missing"
description: "app_build_info has produced no samples for checkout-exporter for at least ten minutes."
runbook: "/posts/prometheus-absent-data-alerts/"
The expression creates one alerting series when the metric is absent, so the for clause can keep a short interruption from opening a ticket. The ten-minute range and two-minute pending period are independent controls: the first asks whether any sample existed recently, while the second asks whether the absence persists across rule evaluations.
Do not use a broad selector such as absent_over_time({job="checkout-exporter"}[10m]) as a universal guard. It can match unrelated series and hide a broken metric contract. Name the metric or use a recording rule that represents the exact required signal. Test the query by temporarily stopping the exporter in a non-production environment and checking both the normal and absent results.
A presence rule also needs a retirement process. If a job is intentionally deleted, remove or change the rule in the same change that removes the target. Otherwise a correct decommission becomes a permanent alert. Treat required metrics like API fields: ownership, compatibility, and removal all need a review path.
Separate missing series from a failed scrape
absent_over_time() answers “did this metric produce any sample?” It does not explain why. Pair it with target health so the first notification points toward the likely cause.
For ordinary Prometheus scrape targets, up is 1 when the last scrape succeeded and 0 when it failed. A simple target-health alert is:
- alert: CheckoutExporterScrapeFailing
expr: up{job="checkout-exporter"} == 0
for: 5m
labels:
severity: page
team: checkout
annotations:
summary: "Checkout exporter scrape is failing"
description: "Prometheus has failed to scrape the checkout exporter for five minutes."
This detects a target that is still discovered but cannot be scraped. It will not catch a target that vanished from service discovery, so keep a separate inventory or discovery-health check when the target set is expected to be stable. The Prometheus scrape configuration defines discovery and relabeling as part of the scrape job; a relabeling change can remove a target before up is available for it.
Use a recording rule to make the distinction visible in dashboards and incident queries:
- record: telemetry:checkout_exporter:target_down
expr: up{job="checkout-exporter"} == 0
- record: telemetry:checkout_exporter:metric_missing
expr: absent_over_time(app_build_info{job="checkout-exporter"}[10m])
The two records can be true at once. That is useful: a down target explains why the required metric is missing. In Alertmanager, group them by job and team rather than sending two unrelated notifications. Alertmanager groups similar alerts, routes them to receivers, and supports inhibition when one root-cause alert should suppress symptoms.
A data-source error is a third case. If the PromQL query cannot be evaluated because a remote store times out or Prometheus is overloaded, absence is not the same as “the metric has no samples.” Keep rule-evaluation and query-error monitoring separate from metric-presence rules. Grafana’s No Data and Error state guidance makes the same distinction: an evaluation that succeeds with no points is different from an evaluation that fails.
Protect multi-dimensional alerts from disappearing labels
A query can return some series while silently losing one region, cluster, or tenant. An expression such as this one only evaluates the regions that still exist:
rate(http_requests_total{service="checkout"}[5m])
If region="us-east" disappears, the remaining regions still produce results. A threshold alert on error rate may remain green because it has no instance for the missing region.
When the expected label set is known, compare telemetry with an inventory metric. For example, suppose service discovery exports one checkout_region_info series for each deployed region:
count by (region) (
checkout_region_info{service="checkout"}
)
unless
count by (region) (
rate(http_requests_total{service="checkout"}[5m])
)
The result is the set of inventory regions with no request series. The exact inventory metric depends on your deployment system, but the pattern is more honest than guessing that an empty result means zero traffic. If a region can be intentionally drained, label that state and exclude it from the expected set through a recording rule or maintenance window.
For a fixed small set, explicit checks are easier to read and test:
absent_over_time(
http_requests_total{service="checkout",region="us-east"}[10m]
)
Avoid creating a high-cardinality alert for every ephemeral Pod. Alert on a stable service, cluster, or region identity, and send Pod-level detail to the dashboard or runbook. The responder needs to know whether the missing series is an isolated restart or a coverage gap across an entire failure domain.
Do not confuse zero work with missing telemetry
The most common no-data mistake is alerting on an ephemeral worker metric. A worker may stop emitting queue_depth because it has scaled down, while the queue exporter continues to expose a real 0. Prefer the exporter or control-plane metric that survives the worker lifecycle.
For Kubernetes autoscaling, the resource metrics pipeline explains that the Metrics API supplies basic CPU and memory data for autoscaling, while custom and external pipelines provide other signals. A queue metric used to wake a worker from zero must remain queryable when no worker Pods run. If it disappears at zero, the HPA cannot distinguish “no work” from “no signal.”
A practical queue policy looks like this:
queue_depth{queue="checkout"} >= 1
That is a demand alert only if queue_depth has a real zero sample when idle. Add a separate exporter-presence rule if the queue metric itself can disappear:
absent_over_time(queue_depth{queue="checkout"}[10m])
The first expression says work exists. The second says the system stopped telling you whether work exists. They should have different severities and different owners. A demand alert may wake the worker platform; a missing-signal alert may open a ticket for the observability owner.
The same principle applies to SLOs. A request-based SLO with no requests is not automatically healthy or unhealthy. You need a denominator policy: synthetic traffic, a longer window, or an explicit low-volume state. Google’s SRE Workbook guidance on alerting SLOs evaluates alert quality using precision, recall, detection time, and reset time. Those criteria are useful for no-data rules too: a rule that pages on every quiet period has poor precision, while a rule that ignores a broken telemetry path has poor recall.
Test the policy before routing it to a human
Treat no-data rules as production code. Prometheus ships promtool for syntax checking; its rule documentation shows how to validate rule files before loading them. Syntax validation is necessary, but it does not prove that a metric disappears in the way your query expects.
Use a small test matrix:
- Normal reporting: required metric exists and
upis1; no alert fires. - One failed scrape:
upis0briefly; only the pending state is entered. - Persistent failed scrape:
upremains0; the target-health alert fires. - Target removed intentionally: the inventory is updated and no stale absence alert remains.
- Metric contract break: the exporter is reachable, but the required metric is absent; the presence alert fires.
- Idle worker: the queue exporter returns
0; no missing-data alert fires and the scale-down remains quiet. - Partial label loss: one region disappears while another reports; the inventory comparison identifies the missing region.
- Query error: the data source times out; an evaluation-error path fires instead of pretending that the metric is empty.
Run each case against the same rule files used in production. Add a runbook link that answers three questions: what signal disappeared, which owner receives the alert, and how to decide whether the target was intentionally removed. If the first action is always “open the dashboard,” the alert is not yet specific enough.
Before enabling a page, simulate the notification path with a non-paging receiver. Confirm that grouping and inhibition collapse a down-target alert plus its missing-metric symptom into one useful message. Then review the alert after the next planned deployment and scale-down. A no-data policy is complete only when normal lifecycle events are quiet and unexpected telemetry loss is visible.
Frequently asked questions
Frequently asked questions
Does Prometheus alert when a query returns no data?+
When should I use absent_over_time() instead of absent()?+
How do I tell a failed scrape from a retired target?+
Should a missing queue metric page the on-call engineer?+
Can a no-data alert detect one missing region?+
Where should no-data alerts be routed?+
Read next
- Set up SLO burn-rate alerts with PromQL
- Check HTTP responses with the HTTP status reference
- Scale a Kubernetes queue worker to zero safely
Tags: #sre, #prometheus, #promql, #alerting, #observability