Why Queues Exist
Every request handler has a budget. A web request that must touch a database, call a third-party API, and wait for the result has a deadline measured in milliseconds; a user sitting in front of a browser has a patience measured in a few seconds. The moment a unit of work stops fitting inside that window — an email that must reach an SMTP server, an image that must be resized, a report that must be generated — you have two choices: make the caller wait for it, or hand it to something that can take its time.
A queue is that "something." It decouples the producer (the code that wants work done) from the consumer (the code that does the work) with an intermediate store. The producer enqueues a message and moves on. The consumer picks messages up, processes them, and acknowledges them. Because producer and consumer no longer share a lifecycle, you get three properties that are hard to get any other way:
- Durability. Work that has been enqueued survives a crash of the process that created it. The request that created a job and then died has still gotten its work done, eventually.
- Backpressure absorption. A spike of 10,000 signups produces a spike of 10,000 jobs, but your workers keep processing at their own rate. The queue absorbs the burst instead of propagating latency to users.
- Isolation. A worker that runs out of memory does not take down the API. A job that fails does not fail the request that created it. Failure is contained inside the background pipeline.
When should you reach for a queue? The useful rule of thumb is any work that is (a) slow, (b) not needed synchronously by the caller, or (c) prone to transient failure. Webhooks to be delivered, files to be processed, notifications to be sent, retried API calls, scheduled maintenance — all natural background jobs. Work that is fast and must block the response should just stay in the request.
But a queue is a sharp tool. It changes your failure model. When the caller used to get an error the moment work failed, now the worker fails in the background, and nobody is watching unless watching is part of the design. Resilience is the discipline that closes that gap.
Delivery Semantics: What "Exactly Once" Really Means
Before writing any worker code, decide what delivery guarantee your system needs, because both the queue and your job handler inherit it. Three terms dominate the discussion:
- At-most-once. The queue makes a best-effort attempt to deliver each message. If the consumer fails or crashes mid-processing, the message is lost. Cheap, simple, and almost never what you want for business-critical work.
- At-least-once. Each message is retried until the consumer acknowledges it. If the consumer crashes after doing the work but before acknowledging, the message is delivered again. This is the default of most serious queue systems — including BullMQ and SQS — and it is the honest guarantee you should design for.
- Exactly-once. Each message is delivered and processed precisely once. This is the guarantee people ask for and the one nobody actually delivers, because it is not achievable in a distributed system subject to network partitions. What systems label "exactly-once" is really at-least-once plus deduplication: a duplicate may arrive, but the system makes the duplicate harmless.
The Honest Guarantee
Treat any claim of exactly-once with suspicion. What you can build — and what this article is really about — is effective exactly-once: at-least-once delivery, idempotent processing, and retries that terminate in a dead-letter queue. If your worker can survive being asked to do the same job twice, the delivery semantics stop being the thing that keeps you up at night.
Idempotency Is the Survival Strategy
An operation is idempotent if performing it twice has the same observable effect as performing it once. It is the single most important property a background job handler can have, because at-least-once delivery guarantees that duplicates will happen. They happen when a worker crashes after committing its work but before acknowledging the job. They happen when two workers pick up the same job during a stall. They happen when you retry a job that actually succeeded.
The first line of defense is making the work itself idempotent. If a job's effect is "set this value," "delete this row," or "charge this user only if no charge exists for this order," the job can simply run twice and the second run is a no-op. Writing idempotent business logic is a design discipline, not a library call.
The second line of defense is a deduplication record. For a job that produces a database effect, store a marker that proves the work is done. Start with a shared Redis connection that every snippet in this article reuses:
import IORedis from 'ioredis';
export const redisConnection = new IORedis({
host: process.env.REDIS_HOST ?? 'localhost',
port: Number(process.env.REDIS_PORT ?? 6379),
maxRetriesPerRequest: null, // BullMQ requires this for worker connections
});Now the worker and the dedup table:
CREATE TABLE processed_jobs (
job_id TEXT PRIMARY KEY,
processed_at TIMESTAMPTZ NOT NULL DEFAULT now(),
result JSONB
);import { Worker } from 'bullmq';
import { pool } from './db'; // a node-postgres Pool
const worker = new Worker('payments', async (job) => {
const claimed = await claimJob(job.id);
if (!claimed) {
// Duplicate delivery: the first run already did the work.
await job.log('duplicate delivery, skipping');
return;
}
await chargeCustomer(job.data.customerId, job.data.amount);
await markProcessed(job.id, { status: 'charged' });
}, { connection: redisConnection });
async function claimJob(jobId: string): Promise<boolean> {
const { rows } = await pool.query(
`INSERT INTO processed_jobs (job_id)
VALUES ($1)
ON CONFLICT (job_id) DO NOTHING
RETURNING job_id`,
[jobId],
);
return rows.length > 0; // false when the job_id already existed
}The ON CONFLICT DO NOTHING pattern is the heart of it: the database — not your application — decides who wins the race when two runs of the same job arrive simultaneously. This is why the unique constraint lives in the database rather than in a client-side check-then-write; the constraint is atomic, and a client check is not. One caveat: a crash between the work and markProcessed means the retry skips correctly only if the work actually happened. Pairing the marker with an idempotent underlying operation covers that gap, which is why the two defenses belong together.
Dedup at Enqueue Time
BullMQ also gives you a lighter-weight deduplication lever at enqueue time. If you give a job a fixed jobId, BullMQ will refuse to enqueue a second job with the same id while the first still exists in the queue:
import { Queue } from 'bullmq';
const paymentsQueue = new Queue('payments', { connection: redisConnection });
await paymentsQueue.add('charge', { customerId: 'cus_42', amount: 4990 }, {
jobId: `charge:${orderId}`,
removeOnComplete: 100, // keep the last 100 completed jobs for inspection
removeOnFail: 1000,
});This stops duplicates at the door instead of at the database — but note it only deduplicates while the original job is still present. A completed job is eventually removed, and a repeated business event could enqueue again. Keep the database-level idempotency guard as the source of truth.
Retries, Backoff, and the Road to the Dead-Letter Queue
Transient failures are the norm, not the exception: a database mid-failover, a rate-limited third-party API, a network blip. The cheapest resilience you can buy is to simply run the job again — with a pause.
BullMQ makes this declarative. Configure attempts and backoff at enqueue time:
await paymentsQueue.add('charge', payload, {
attempts: 5,
backoff: {
type: 'exponential',
delay: 2000, // base delay in milliseconds
},
removeOnComplete: 100,
removeOnFail: 1000,
});With attempts: 5 and exponential backoff starting at 2 seconds, the job waits roughly 2s, 4s, 8s, and 16s between retries — about 30 seconds of cooling before the job gives up. Exponential backoff matters because the alternative, fixed-interval retries, lets a downstream outage receive your retries in lockstep, all at once, at the worst possible moment. Exponential backoff spreads those retries out and gives the failing dependency room to recover.
Two Kinds of Failure
A job can fail because the system is temporarily broken (retry: yes) or because the input is permanently wrong (retry: pointless — a malformed payload will fail again in exactly the same way). Separate the two with distinct error classes, and let permanent failures skip the retry loop entirely. BullMQ supports this through a custom backoffStrategy, which receives the error and can return a negative value to stop retrying:
import { Worker } from 'bullmq';
class PermanentlyFailingError extends Error {}
const worker = new Worker('payments', processCharge, {
connection: redisConnection,
backoffStrategy: (attemptsMade: number, _type?: string, err?: Error) => {
if (err instanceof PermanentlyFailingError) return -1; // do not retry
return 2 ** attemptsMade * 1000; // exponential
},
});The retry budget itself is a business decision. A webhook that must eventually reach a customer's system deserves more attempts than a cache warm that is harmless to lose.
The Dead-Letter Queue
When attempts are exhausted, the job has nowhere left to go in the main queue — and if you do nothing, it silently sits in the failed set and becomes invisible. That is the moment a dead-letter queue (DLQ) enters. A DLQ is a second queue whose only job is to catch work that permanently failed, so a human or an alerting system can look at it. It is the difference between "a job failed" and "nobody noticed."
import { Queue, QueueEvents } from 'bullmq';
const paymentsQueue = new Queue('payments', { connection: redisConnection });
const dlq = new Queue('payments-dlq', { connection: redisConnection });
const events = new QueueEvents('payments', { connection: redisConnection });
events.on('failed', async ({ jobId, failedReason }) => {
const job = await paymentsQueue.getJob(jobId);
if (!job) return;
const gaveUp = job.attemptsMade >= (job.opts.attempts ?? 1);
if (gaveUp) {
await dlq.add('failed-charge', {
originalJobId: job.id,
data: job.data,
failedReason,
attemptsMade: job.attemptsMade,
failedAt: new Date().toISOString(),
});
}
});Move everything needed to diagnose the failure into the DLQ payload — the original data, the failure reason, the attempt count. The DLQ entry is a post-mortem, not just a copy of the job.
Scaling Workers Without Breaking Guarantees
A single worker process is fine until it is not. Scaling in BullMQ means running more workers — more concurrency inside one process, more processes on one machine, or more machines — all consuming from the same queue. Because BullMQ is backed by Redis, the queue itself is shared state, and the workers coordinate through it rather than through each other.
Two knobs matter:
- Concurrency is the number of jobs one worker processes at the same time, set at worker construction:
const worker = new Worker('payments', processCharge, {
connection: redisConnection,
concurrency: 10, // this process runs up to 10 jobs in parallel
});More concurrency means higher throughput per process, but it also means your job handler must be safe under parallel execution. Shared in-memory state between jobs will start corrupting. If the handler touches a shared cache or connection, make it concurrent-safe first.
- Worker count is horizontal: N processes, possibly on N machines, each running the same worker code. The queue distributes jobs among them. Add capacity by adding processes, remove it by scaling down. There is no leader election to manage and no shared worker state to coordinate — that is the point.
But scaling changes your guarantees in two subtle ways. First, at-least-once becomes more visible: every additional worker is another place a job can be fetched and lost mid-processing. The system already tolerates this — that is why idempotency is mandatory — but the rate of duplicate delivery grows with the number of workers. Do not postpone the idempotency work because "we only have one worker today."
Stalls: A Second Source of Duplicates
Second, stalls multiply. BullMQ workers lease jobs from Redis for the duration of processing. A worker that crashes, is killed by OOM, or is deployed under your feet dies mid-job. After a stall-detection interval, BullMQ treats the job as stalled and re-delivers it, controlled by maxStalledCount (default 1). Stalled jobs are a second source of duplicate delivery, entirely independent of your retry policy:
const worker = new Worker('payments', processCharge, {
connection: redisConnection,
maxStalledCount: 2, // tolerate more stalls before giving up on a job
});
worker.on('stalled', (jobId: string) => {
console.error(`job ${jobId} stalled and will be re-delivered`);
});The practical consequence: a horizontal worker fleet only works if every job is idempotent and every job has a bounded retry budget. The more parallelism you add, the more duplicates you can expect, and the more the database-level dedup earns its keep.
Scheduling Work with Repeatable Jobs
Not all background work is triggered by an event. Digest emails, cache refreshes, nightly reindexes, subscription renewals — these run on a schedule. BullMQ folds scheduling into the same primitive: a repeatable job is a job with a repeat rule.
// Run every night at 02:15.
await queue.add('nightly-reindex', { index: 'search' }, {
repeat: { pattern: '15 2 * * *' },
jobId: 'nightly-reindex',
});
// Run every 15 minutes.
await queue.add('cache-refresh', { namespace: 'top-products' }, {
repeat: { every: 15 * 60 * 1000 },
jobId: 'cache-refresh',
});Two things to get right with scheduled jobs. First, use a stable jobId: a repeatable job identified only by its pattern and data will be re-created on every deploy if you re-add it at startup. A stable id turns re-adding it into a no-op instead of a growing stack of duplicate schedules. Second, understand what every means: it schedules the next run from the moment the current run fires, not on a fixed wall-clock grid. For true cron behavior, use the pattern form, which BullMQ evaluates like a standard five-field crontab.
Also schedule a job to check your workers. A recurring heartbeat job that runs every minute, does nothing useful itself, and fails loudly if it never runs is a cheap detector for a queue pipeline that has silently stopped — the monitoring section below relies on exactly this kind of signal.
Monitoring: Queue Depth, Stalls, and Dead-Letter Rate
A queue gives you asynchronous work and, with it, asynchronous failure. If nobody watches, the failures are invisible. Monitoring a queue system comes down to three numbers and the story they tell:
- Queue depth — the number of jobs
waitingplusdelayed. A queue that grows monotonically is a worker deficit, not a mystery. - Stalled rate — how often jobs get re-delivered because a worker died. A rising stalled rate is a worker-health problem: OOM, crashing deploys, code that throws outside the job handler.
- Dead-letter rate — how often work exhausts its retries. A rising DLQ rate is a business-level failure signal: a payment integration that broke, an API contract that changed.
BullMQ exposes these numbers directly:
const counts = await queue.getJobCounts(
'waiting', 'active', 'delayed', 'failed', 'completed',
);
// { waiting: 12, active: 4, delayed: 0, failed: 3, completed: 840 }
const metrics = await queue.getMetrics('failed');Export these as Prometheus gauges with a small exporter loop, and alert on the conditions that indicate actual damage rather than noise:
| Metric | Alert condition | What it usually means |
|---|---|---|
| Queue depth | waiting grows for N consecutive minutes |
Workers under-provisioned or stuck |
| Active jobs | Active near max concurrency for long periods | Backpressure reaching the ceiling |
| Stalled rate | Rises across a deploy window | Worker crashes, OOM, deploy churn |
| Dead-letter rate | DLQ receives new jobs | A dependency or contract broke permanently |
| Age of oldest waiting job | Exceeds a threshold | Head-of-line blocking or a stalled producer |
Alerting on the Right Numbers
Alerting discipline matters as much as metrics. Queue depth is a leading indicator and worth alerting on. Dead-letter growth is a trailing indicator — by the time a DLQ fills, the work has already failed — but it is the signal that a business process broke, so keep its threshold strict. And always pair a metric with a human-readable story: "dead-letter rate above 5 per hour" tells an on-call engineer more than "failed: 42".
Failure Modes and Defenses
The failure modes below are the ones that actually bite queue systems in production. Defense in depth — each row of the table is an independent layer — is the point.
| Failure mode | What happens | Defense |
|---|---|---|
| Worker crash mid-job | Job re-delivered; double side effects | Idempotent handlers; DB dedup |
| Job stalls | Re-delivered after the stall interval | Short jobs; progress heartbeat; watch stalled |
| Retries exhaust | Job invisible in the failed set |
Dead-letter queue; alert on DLQ growth |
| Poison message | A job that always fails consumes its retries | Permanent-error class skips backoff, goes straight to DLQ |
| Producer enqueues duplicates | The queue processes the same event twice | Stable jobId; database unique constraint |
| Worker OOM / deploy churn | Fleet thrashes, stalls spike | Graceful shutdown; stagger deploys |
| Redis unavailable | Enqueue and processing both halt | Redis HA; fail the request loudly, never silently |
| Handler grows unbounded | Memory pressure kills the worker | Concurrency budget; batch; stream results to storage |
Two of these deserve emphasis. Graceful shutdown is a worker hygiene rule: on SIGTERM, stop accepting new jobs, let in-flight jobs finish, and only then exit — otherwise every deploy turns into a stall storm:
import { Worker } from 'bullmq';
const worker = new Worker('payments', processCharge, { connection: redisConnection });
async function shutdown(signal: string) {
console.log(`received ${signal}, closing worker`);
await worker.close(); // stops new jobs, waits for active ones
process.exit(0);
}
process.on('SIGTERM', () => shutdown('SIGTERM'));
process.on('SIGINT', () => shutdown('SIGINT'));And keep jobs short. BullMQ holds a lock on a job for the duration of processing; a job that runs for an hour holds that lock for an hour and, if it dies, contributes to stalled re-delivery of everything it held. If a unit of work naturally takes a long time, split it into smaller jobs and chain them, or report progress along the way. Short jobs are also what make horizontal scaling work: more workers only help if the unit of work they grab is small enough to distribute.
Conclusion
A queue moves work out of the request path and gives it durability, but it moves the failure with it. The resilient design is a stack of independent layers, each of which assumes the ones below it failed:
- At-least-once is your honest delivery guarantee. Design for duplicates, not against them.
- Idempotency is the floor. Database-level dedup, so the second run of any job is a no-op.
- Retries are cheap resilience. Exponential backoff, a bounded attempt budget, and permanent failures that skip the retry loop.
- The DLQ is the last stop. Work that gives up lands somewhere visible, with a post-mortem attached.
- Scaling is safe only because of the layers above it. Add workers freely, but never before idempotency and retry budgets exist.
- Scheduled work uses the same machinery. Repeatable jobs, stable ids, and a heartbeat that proves the pipeline is alive.
- Monitoring closes the loop. Queue depth, stalled rate, and dead-letter rate are the three numbers that tell you whether background work is healthy.
None of these layers is clever. Each is a small, boring mechanism — a unique constraint, a backoff curve, a second queue, a gauge. Put together, they turn a queue from a place where work goes to die into a place where work eventually, reliably, gets done.