Durabull Documentation

Redis Health Alerts

Monitor Redis INFO memory, CPU, client, eviction, and connection-pressure thresholds with durable Durabull incidents.

Redis health alert rules watch the Redis server behind a Durabull connection. They use one lightweight INFO sample per alert-monitor poll, regardless of how many BullMQ queues use the connection.

The same sample feeds the Analytics tab, where memory allocation, resident memory, configured capacity, CPU, client utilization, fragmentation, evictions, blocked clients, and rejected connections are charted over time. History collection is enabled by default and does not require a Redis Health alert rule.

Observe Redis Health Over Time

Open a Redis connection and select Analytics. The Redis Resource Health section shows:

  • Current resource values and whether the sample is fresh.
  • Utilization, memory-footprint, and pressure-signal time series for the selected 1-hour to 30-day window.
  • Active Redis Health alert thresholds overlaid with the telemetry they evaluate.
  • Explicit gaps and a coverage percentage when collection was interrupted or INFO was unavailable.

Coverage uses the configured poll interval as its expected cadence, so a healthy five-minute collector is not reported as 80% missing.

Long windows are downsampled in the database before they are returned to the browser, so the chart response stays bounded even on high-volume installations. Downsampled buckets retain each metric's peak rather than its average, so a short threshold-crossing excursion remains visible.

Create a Redis Health Rule

  1. Open a Redis connection and select Alerts.
  2. Select Create rule.
  3. Choose Redis memory pressure for the 80% memory template, or choose Redis Health under Condition.
  4. Select the Redis INFO metric and the threshold that should open an incident.
  5. Add saved or one-off notification destinations, name the rule, and save it.

Use Run live test from an existing rule to evaluate the latest INFO sample without creating an incident. Email, webhook, and Linear routing uses the same durable delivery and retry pipeline as queue alerts.

Available Metrics

MetricINFO fieldsEvaluation
Memory usageused_memory, maxmemoryPercent of the configured maxmemory; unavailable when Redis has no explicit limit.
Memory allocatedused_memoryAbsolute Redis allocation in MiB.
Resident memoryused_memory_rssPhysical memory occupied by the Redis process in MiB.
CPU usageused_cpu_sys, used_cpu_userCPU seconds consumed between two monitor samples, divided by elapsed wall time.
Allocator fragmentationallocator_frag_ratio, allocator_frag_bytesRatio threshold; overhead below 10 MiB is ignored to avoid noisy high ratios on small instances.
Client capacityconnected_clients, maxclientsConnected clients as a percentage of the configured limit.
Blocked clientsblocked_clientsCurrent blocked-client count.
Key evictionsevicted_keysCounter increase normalized to keys per minute.
Rejected connectionsrejected_connectionsCounter increase normalized to rejected connections per minute.

Thresholds fire when the measured value is equal to or greater than the configured value. An open incident resolves automatically after a later poll observes that the condition has cleared. Missing INFO data never resolves an open incident because an unavailable measurement is not proof of recovery.

CPU, eviction, and rejection rates need two samples. A new rule stays quiet on its first sample and can evaluate on the next poll. The previous sample is stored in Durabull's alert cursor, so rate calculations continue across API restarts. Redis counter resets are treated as zero activity rather than a spike.

Operational Requirements

  • The Durabull API and its background monitor must remain running. DURABULL_ALERT_POLL_INTERVAL_MS controls the sampling interval for both alert evaluation and Redis history; DURABULL_ALERT_ENABLED independently controls alert evaluation.
  • Redis history collection is controlled independently with DURABULL_REDIS_HEALTH_HISTORY_ENABLED (default true). Collection continues when alerting is disabled so Analytics does not lose its server-health timeline.
  • Durabull stores at most one health sample per minute per connection. A cleanup job deletes samples older than DURABULL_REDIS_HEALTH_RETENTION_DAYS (default 30, bounded to the observable 1–30 day window), even when new history collection is disabled. Cleanup deletes in 5,000-row transactions with a 100,000-row per-run safety limit. If it finds a larger backlog, it schedules another bounded catch-up run after five seconds instead of waiting for the next hourly sweep or creating one large blocking transaction.
  • The Redis user must be allowed to run INFO. Managed services can restrict the command or omit fields; a rule stays quiet when its selected metric is unavailable.
  • When maxmemory is zero, percentage-based memory usage is unavailable because host memory can overstate the capacity available to Redis in containers or managed services. Use an absolute allocation or resident-memory rule, or configure an explicit Redis maxmemory limit.
  • A CPU result can exceed 100% because Redis reports CPU across the server process's threads.

Redis INFO collection failures are isolated from BullMQ queue checks: a restricted or incomplete INFO response does not stop existing queue alert rules from being evaluated.