Introduction
Background jobs run without user-facing feedback, making monitoring essential. Without visibility into queue depth, latency, and failure rates, problems go undetected.
Key Concepts
- Queue Depth: Jobs waiting to be processed.
- Queue Latency: Time between enqueue and execution start.
- Failure Rate: Percentage of jobs that fail.
Real World Context
At 2 AM, your payment gateway starts timing out. Without monitoring, failed jobs pile up silently. With monitoring, an alert fires immediately.
Deep Dive
Solid Queue Monitoring
rubySolidQueue::ReadyExecution.count SolidQueue::ScheduledExecution.count SolidQueue::FailedExecution.count SolidQueue::ReadyExecution.where(queue_name: "default").count
Health Check Endpoint
rubyclass HealthController < ApplicationController def jobs ready = SolidQueue::ReadyExecution.count failed = SolidQueue::FailedExecution.count healthy = failed < 100 && ready < 10_000 render json: { healthy: healthy, ready: ready, failed: failed }, status: healthy ? :ok : :service_unavailable end end
ActiveSupport Notifications
rubyActiveSupport::Notifications.subscribe("perform.active_job") do |*args| event = ActiveSupport::Notifications::Event.new(*args) job = event.payload[:job] Rails.logger.info "Job: #{job.class.name} queue=#{job.queue_name} duration=#{event.duration.round(1)}ms" end
Common Pitfalls
- Only monitoring failures — A queue with zero failures but 10-minute latency is still broken.
- Alert thresholds too low — Normal spikes (deploys) should not page your team.
Best Practices
- Create a /health/jobs endpoint for external monitoring.
- Subscribe to Active Job notifications for standardized metrics.
Summary
- Monitor queue depth, latency, and failure rates.
- Query SolidQueue tables or use ActiveSupport notifications.
- Create health check endpoints for external monitoring.
- Alert on sustained high latency and rising failure counts.
Code Examples
ruby
ActiveSupport::Notifications.subscribe("perform.active_job") do |*args|
event = ActiveSupport::Notifications::Event.new(*args)
job = event.payload[:job]
StatsD.timing("jobs.duration", event.duration, tags: ["class:#{job.class.name}"])
end