Your cart is currently empty!
Scheduling and retrying a scrape without making things worse
I once wrote a postmortem blaming a supplier for flaky proxy lines. I had to send a second one a week later saying the lines were fine and the problem was me.
The job config said three retries. The http client underneath it was also set to three. Nine attempts per failing page, forty workers wide, all triggered by one five minute wobble on the target. About three hundred and sixty requests hit that site inside a few seconds, and it did what any sane infrastructure does.
The scraper worked perfectly. The layer above the scraper is what broke, and that layer is what this piece is about.
The number in your config is not the number
Retry counts multiply across layers, and almost nobody writes down the product.
Scrapy retries twice by default through RETRY_TIMES, so a fresh project is already at three attempts before you touch anything. Most cloud sdks retry internally before your code ever sees a failure. Wrap any of that in an orchestrator with its own task retry setting and you have built a multiplier without noticing.
The fix is a decision, not a library. Exactly one layer owns retries, everything below it fails fast and reports upward, and the effective number lives in a comment where a human will read it. For me the owner is the job runner, because it is the only thing that can see how many workers are already in flight. The http client sits at zero and says zero out loud.
What an immediate retry actually does
A retry is a purchase. You are spending a second request against a target that just told you it could not serve the first one, and the naive version spends it at the worst possible moment.
Think about what the target sees. Forty workers fail at the same instant because the cause was shared. All forty come straight back. Same paths, same second, same address ranges. Any site with rate defences reads that as a signature and acts on it.
This is the most common way people get themselves blocked while filing a ticket about the target.
Backoff, and where to cap it
Exponential backoff is one line of arithmetic: wait base * 2 ** attempt. One second, two, four, eight. It works because the delay grows faster than most transient problems last, so a blip costs you a second and a genuine outage parks you somewhere sensible.
Two things people get wrong with it.
Cap the delay. Unbounded doubling will eventually sleep a worker for half an hour, and a worker asleep for half an hour is capacity you are paying for and not using. Sixty seconds is a fine ceiling for most scrapes.
And honor an explicit instruction over your own formula. A 429 usually arrives with a Retry-After header carrying the exact number of seconds the site wants. Your curve is a guess. That header is not. Coming back sooner than a site asked in writing is the rudest thing in the whole stack, and it is the fastest route from a temporary throttle to a durable block.
Jitter is what stops a fleet from reconverging
Backoff alone spreads nothing out.
Forty workers fail in the same second, all compute the same two second wait, and all return in the same second. You moved the spike two seconds to the right. Its shape is identical.
Jitter means randomising the wait: random() * base * 2 ** attempt instead of the bare formula. Now those forty land smeared across a window.
It matters more the more machines you own. I have scrape jobs on four boxes. If all four use the same formula against the same clock, they do not collide once and recover. They reconverge. They fail together, back off by the same amount, and find each other again on every round after that, which is how a fleet of well behaved schedulers produces a synchronised spike that none of them individually asked for.
One question, instead of a table of status codes
Before retrying anything, ask whether anything will be different next time.
If nothing about the request or the world has changed, the retry is asking the same question louder. A permission error is the clean example: credentials that were wrong a second ago are still wrong, and five more attempts produce five identical refusals plus a pattern in someone’s logs.
That question sorts nearly everything, and it keeps working on errors you have never seen before.
Worth retrying, because the cause sits outside your request and is temporary: connection timeouts, resets, 5xx from the origin, and a 429 where you actually wait as instructed.
Not worth retrying: 401 and 403, parse failures, validation failures. A 403 is a decision the site made about the identity behind the request, so repeating the identical request from that identical identity mostly confirms the decision. A parse error retried five times gives you five copies of one stack trace.
404 sits between the two. A URL that worked last week and 404s today is genuinely gone, so log it once and mark it resolved rather than retrying it. A URL you invented by incrementing an id was never there, and that is not a failure at all.
Attempt ceilings and the pile behind them
Every job needs a hard cap. Without one a retry loop is an infinite loop with better manners.
When an item hits the cap it goes somewhere durable with the error attached and the input that produced it. A dead letter table with the url, the attempt count, the last error and a timestamp is enough. So is a folder of json files.
The ceiling on its own is half a solution. I have opened dead letter tables holding forty thousand rows, months of quiet accumulation, with no alert ever fired because technically nothing failed. The items were caught, which is exactly what made them invisible.
So alert on the pile as a rate, not a count. Total size only tells you the job is old. New rows per day compared against a normal week tells you something started breaking on Tuesday.
Idempotency is the permission slip
All of this rests on the job being safe to run twice.
Idempotent means a second run leaves the database exactly where the first one did: same rows, same counts, nothing doubled. With that property a retry is a shrug and a backfill is a command. Without it, every retry is a gamble.
You feel the difference at two in the morning, in the moment you hesitate before clicking rerun. That hesitation is the real cost of a job you cannot safely repeat.
A checkpoint turns a dead job into a resumed one
A long job that dies at eighty percent and restarts from zero has charged you that eighty percent twice, and charged the target for it too.
Checkpoint the cursor. Write down the last page or id you finished, durably, after the rows commit and never before. Do it in the same transaction as the write if your storage allows: a checkpoint that lands before the data produces silent gaps rather than duplicates, and gaps are far harder to notice.
This is the cheapest item on the list and the one skipped most often, because for the first month the job never dies.
Nobody needs to start at zero minutes past
Open your crontab and count how many entries start at minute zero.
So does everyone else’s. Every hourly scrape, plus the backups, the reports and the health checks, all firing on the hour because nobody chose otherwise. Your target meets the worst minute of its day sixty times a day at the same offset, and you are standing in that crowd for no benefit.
Move off it. Pick 17 past, or hash the job name into a minute so the offset is automatic and stable across deploys: minute = crc32(job_name) % 60. The request is identical, the data is identical, you have just stopped arriving with the mob.
It helps your own side too. I had four jobs on the same box all set to minute zero, fighting over one proxy pool and one cpu for four minutes an hour and idling for the other fifty six. They sit at 7, 19, 31 and 43 now, and the box is boring.
Alert on the shape of the output
An exit code of zero means your code reached the end without throwing. That is the entire claim it makes.
A scraper that fetched a block page, parsed nothing out of it, wrote zero rows and returned cleanly is a success by that measure. Every dashboard is green and the data is gone.
So make the run’s own output the alert. Count what you produced, compare against this source’s recent normal, and fail the run deliberately when the number is absurd. A source that gives you eleven thousand rows most days and six today should be red even though nothing errored. Zero rows is the loud version. One field going null while the rows still validate is the quiet one. A run that never started hides longest, because there is no run to inspect and no error to catch, so the absence of a successful run needs its own alert.
Left to exit codes, these get found when somebody downstream asks why a number looks stale. That is a detection time measured in weeks.
What a retry policy does not buy you
None of this changes what you are allowed to collect. Backing off is restraint, and restraint does not widen the target by an inch. Public data, a robots file honored, a published crawl delay treated as an instruction, and an official api or bulk feed used whenever one exists.
I run these scrapers over my own carrier lines every day, so the numbers here are ones I have paid for. The worked versions, the backoff and jitter code, the checkpoint pattern and the alert thresholds I actually use are at dataresearchtools.com.
Get new guides and videos first — join the Telegram channel.
Leave a Reply