Skip to content

Alerting

Mimir evaluates the rules and sends the alerts: its ruler queries the metrics every minute, and its built-in Alertmanager groups, routes and delivers what fires. Grafana is not involved, so the rules are plain Prometheus rule files and work unchanged on anything Prometheus-compatible.

Piece File Mimir flags
Rules flux/50_apps/observability/mimir/rules.yaml -ruler-storage.backend=local, -ruler-storage.local.directory=/etc/mimir/rules
Routing and receivers flux/50_apps/observability/mimir/alertmanager.yaml -target=all,alertmanager, -alertmanager-storage.backend=local, -alertmanager-storage.local.path=/etc/mimir/alertmanager

Both are ConfigMaps mounted read-only, and Mimir re-reads them on its own poll, so a change lands without a restart. The local backends cannot be written through the API, which keeps git the only source.

Rules belong to a tenant

The ruler evaluates each tenant's rules against that tenant's metrics only. The local store reads <directory>/<tenant>/<file>.yaml, so the CBC rules ConfigMap is mounted at /etc/mimir/rules/CBC. A rule under any other directory name queries a tenant that holds none of these metrics and silently never fires. Postgres, like everything outside the ABC namespaces, is ingested as CBC (see Tenancy). Rules for another tenant need their own ConfigMap and mount.

Alertmanager is per tenant too: CBC.yaml is the configuration for tenant CBC, the file name being the tenant ID.

Routing by channel

A rule's channel label decides whether it notifies, and where. Each value has its own route and receiver:

routes:
  - matchers: ['channel="infra"']
    receiver: infra
  - matchers: ['channel="backup"']
    receiver: backup

A rule without a channel lands in silent: it still fires and shows in Mimir and Grafana, but never notifies. That keeps notifications to the few rules that need someone now; promoting a rule is adding the label.

Every receiver is empty today, so nothing notifies yet. Mattermost incoming webhooks accept Slack's format, so a channel becomes slack_configs with the webhook URL. To keep that URL out of git, the Alertmanager configuration then moves from this ConfigMap to a Secret rendered by an ExternalSecret template from OpenBao, mounted at the same path.

The Postgres rules

Alert Channel Fires when
PostgresBackupTooOld backup no successful base backup for 26 hours
PostgresBackupFailed backup the latest base backup attempt failed
PostgresBackupMetricsMissing backup the Barman metrics are absent, which would keep the age alert from ever firing
LastFailedArchiveTime backup the primary's latest WAL archive attempt failed
PostgresInstanceDown infra an instance reports Postgres down
PostgresMetricsMissing infra no instance in postgres reports metrics at all
CloudNativePG defaults — long transactions, waiting backends, XID age, replication lag, deadlocks, a failing replica

The cnpg-default group is CloudNativePG's recommended rule set from its 1.30 sample prometheusrule.yaml, minus its archiving rule, which moved to backup. The backup metrics come from the Barman Cloud plugin (barman_cloud_cloudnative_pg_io_*); WAL archiving uses CloudNativePG's own cnpg_pg_stat_archiver_*, since the plugin exports nothing for it. See Backups.

Seeing alerts in Grafana

Mimir's ruler and Alertmanager APIs serve one tenant per request; a federated X-Scope-OrgID such as CBC's ABC-CRM|…|CBC gets "no valid org id found". The CBC org therefore has two datasources pinned to CBC alone, defined in terraform/60_grafana-config/main.tf:

Datasource Shows
Mimir (CBC only, alerting) the rule groups and their state
Mimir Alertmanager (CBC) firing alerts and silences

The federated Mimir datasource has manageAlerts off, so Grafana does not try the ruler through it. Queries and dashboards keep using the federated one.

Every Loki datasource has manageAlerts off as well. No Loki rules exist, and the federated CBC one would fail on the ruler API the same way.