Redis Health Alerts
Monitor Redis INFO memory, CPU, client, eviction, and connection-pressure thresholds with durable Durabull incidents.
Redis health alert rules watch the Redis server behind a Durabull connection. They use one
lightweight INFO sample per alert-monitor poll, regardless of how many BullMQ queues use the
connection.
The same sample feeds the Analytics tab, where memory allocation, resident memory, configured capacity, CPU, client utilization, fragmentation, evictions, blocked clients, and rejected connections are charted over time. History collection is enabled by default and does not require a Redis Health alert rule.
Observe Redis Health Over Time
Open a Redis connection and select Analytics. The Redis Resource Health section shows:
- Current resource values and whether the sample is fresh.
- Utilization, memory-footprint, and pressure-signal time series for the selected 1-hour to 30-day window.
- Active Redis Health alert thresholds overlaid with the telemetry they evaluate.
- Explicit gaps and a coverage percentage when collection was interrupted or
INFOwas unavailable.
Coverage uses the configured poll interval as its expected cadence, so a healthy five-minute collector is not reported as 80% missing.
Long windows are downsampled in the database before they are returned to the browser, so the chart response stays bounded even on high-volume installations. Downsampled buckets retain each metric's peak rather than its average, so a short threshold-crossing excursion remains visible.
Create a Redis Health Rule
- Open a Redis connection and select Alerts.
- Select Create rule.
- Choose Redis memory pressure for the 80% memory template, or choose Redis Health under Condition.
- Select the Redis INFO metric and the threshold that should open an incident.
- Add saved or one-off notification destinations, name the rule, and save it.
Use Run live test from an existing rule to evaluate the latest INFO sample without creating an incident. Email, webhook, and Linear routing uses the same durable delivery and retry pipeline as queue alerts.
Available Metrics
| Metric | INFO fields | Evaluation |
|---|---|---|
| Memory usage | used_memory, maxmemory | Percent of the configured maxmemory; unavailable when Redis has no explicit limit. |
| Memory allocated | used_memory | Absolute Redis allocation in MiB. |
| Resident memory | used_memory_rss | Physical memory occupied by the Redis process in MiB. |
| CPU usage | used_cpu_sys, used_cpu_user | CPU seconds consumed between two monitor samples, divided by elapsed wall time. |
| Allocator fragmentation | allocator_frag_ratio, allocator_frag_bytes | Ratio threshold; overhead below 10 MiB is ignored to avoid noisy high ratios on small instances. |
| Client capacity | connected_clients, maxclients | Connected clients as a percentage of the configured limit. |
| Blocked clients | blocked_clients | Current blocked-client count. |
| Key evictions | evicted_keys | Counter increase normalized to keys per minute. |
| Rejected connections | rejected_connections | Counter increase normalized to rejected connections per minute. |
Thresholds fire when the measured value is equal to or greater than the configured value. An open incident resolves automatically after a later poll observes that the condition has cleared. Missing INFO data never resolves an open incident because an unavailable measurement is not proof of recovery.
CPU, eviction, and rejection rates need two samples. A new rule stays quiet on its first sample and can evaluate on the next poll. The previous sample is stored in Durabull's alert cursor, so rate calculations continue across API restarts. Redis counter resets are treated as zero activity rather than a spike.
Operational Requirements
- The Durabull API and its background monitor must remain running.
DURABULL_ALERT_POLL_INTERVAL_MScontrols the sampling interval for both alert evaluation and Redis history;DURABULL_ALERT_ENABLEDindependently controls alert evaluation. - Redis history collection is controlled independently with
DURABULL_REDIS_HEALTH_HISTORY_ENABLED(defaulttrue). Collection continues when alerting is disabled so Analytics does not lose its server-health timeline. - Durabull stores at most one health sample per minute per connection. A cleanup job deletes samples
older than
DURABULL_REDIS_HEALTH_RETENTION_DAYS(default30, bounded to the observable 1–30 day window), even when new history collection is disabled. Cleanup deletes in 5,000-row transactions with a 100,000-row per-run safety limit. If it finds a larger backlog, it schedules another bounded catch-up run after five seconds instead of waiting for the next hourly sweep or creating one large blocking transaction. - The Redis user must be allowed to run
INFO. Managed services can restrict the command or omit fields; a rule stays quiet when its selected metric is unavailable. - When
maxmemoryis zero, percentage-based memory usage is unavailable because host memory can overstate the capacity available to Redis in containers or managed services. Use an absolute allocation or resident-memory rule, or configure an explicit Redismaxmemorylimit. - A CPU result can exceed 100% because Redis reports CPU across the server process's threads.
Redis INFO collection failures are isolated from BullMQ queue checks: a restricted or incomplete INFO response does not stop existing queue alert rules from being evaluated.