Jobs / Retries and backoff
Retries and backoff
A job that fails is not finished. Larkspur puts it back on the queue with a delay that grows each time, so a flaky dependency gets room to recover instead of a stampede.
How a retry is scheduled
When a handler throws, or returns a value marked Retry, the worker does not drop the job. It writes the error to the job's history, increments attempt, and computes the next run time from the job's retry policy. The job then waits in the scheduled set until that time passes.
Only the worker that held the lease can reschedule a job. If the worker dies mid-run, the lease expires after visibility_timeout and another worker picks the job up as if it had failed. That attempt counts.
What counts as a failure
- An uncaught exception from the handler.
- A return value of
Retry(after=...), which also overrides the delay once. - A timeout, when the handler runs past
timeout.
Returning Discard ends the job without a retry. Use it for input that will never succeed, such as a deleted account.
Backoff with jitter
The default policy is exponential: each delay is the base multiplied by two to the power of the attempt, capped at a ceiling. Without jitter, a thousand jobs that failed together retry together, and the dependency that fell over falls over again. Larkspur applies full jitter by default, picking a random delay between zero and the computed value.
# larkspur.toml
[retry]
policy = "exponential"
base = "2s"
ceiling = "15m"
max_tries = 8
jitter = "full"
The defaults, written out. Omit any key to keep its default.
With these values, attempt four waits up to 32 seconds and attempt eight waits up to the 15 minute ceiling. The total worst case across eight tries is a little over 25 minutes.
| Key | Default | Meaning |
|---|---|---|
base | 2s | Delay before the first retry, before jitter. |
ceiling | 15m | No single delay is longer than this. |
max_tries | 8 | Total attempts, including the first run. |
jitter | full | full, equal, or none. |
Idempotency keys
A retry runs the handler again from the top. If the handler charged a card before it crashed, the retry charges it twice. Give each job an idempotency key and check it before any side effect.
@job(retry="payments")
def capture(ctx, invoice_id):
key = f"capture:{invoice_id}"
if ctx.seen(key):
return Discard
gateway.capture(invoice_id, idempotency_key=key)
ctx.mark(key, ttl="7d")
ctx.seen reads from the same store as the queue. Do not keep keys in process memory; a retry often lands on a different worker.When retries run out
After max_tries, the job moves to the dead-letter queue with its full error history. Nothing runs it again on its own. Inspect it with larkspur dead list, fix the cause, then replay it with larkspur dead retry <id>.
Dead jobs are kept for 14 days by default. Set dead_ttl in the [queue] table to change this.
Alerting on dead jobs
Larkspur emits larkspur.job.dead once per job. Alert on its rate, not on single events. One dead job a day is normal for most queues; forty in five minutes is an outage.
Limits
A policy can set max_tries up to 50 and ceiling up to 24 hours. Longer waits belong in a scheduled job, not a retry. Retries do not survive a queue rename; jobs in the scheduled set keep their old queue name until they run.