Idempotent scraper design: why reruns shouldn’t double your rows

The 2am page that taught me this

A few years back I had a scraping job that pulled product listings every six hours. The orchestrator’s health check timed out at 30 seconds. The page, on a bad network night, took 34 seconds to fully render and get scraped. The orchestrator declared the job dead and kicked off a retry while the original process was still mid-write. Two processes, same target, same time window, both inserting rows into the same table.

By morning we had roughly double the expected row count for that run. Nothing crashed, nothing errored, no alert fired. The pipeline looked healthy. The only sign something was wrong was a downstream dashboard where “unique products” jumped 40% overnight with no matching jump in the actual catalog.

That’s the failure mode idempotent scraper design exists to prevent. Not the scraper crashing, but the scraper succeeding twice.

What “idempotent” actually means for a scraper

In plumbing terms, an idempotent operation produces the same end state no matter how many times you run it. Run it once, run it five times, the database looks identical. For a scraper this means: if the job reruns because of a retry, a cron overlap, a manual backfill, or an operator rerunning a failed batch, the table doesn’t grow just because the code executed again.

This is different from the job being correct. A scraper can extract data perfectly and still not be idempotent, because idempotency is a property of the write path, not the extraction logic. The parsing can be flawless and you can still double your rows if the insert logic doesn’t know how to recognize “I’ve already written this.”

Natural keys beat auto-increment IDs

The single biggest source of duplicate rows I’ve seen in production scrapers is relying on an auto-increment primary key with no unique constraint on anything else. Every insert gets a new ID, so nothing in the database structure stops the same listing from being written twice.

The fix is to identify what actually makes a row unique in the source system and enforce that at the schema level. For a product scrape that might be (marketplace, product_id). For a job listing scrape it might be (source_url, posted_date) if the site reuses URLs for reposted listings. For anything without a stable ID from the source, you fall back to a composite of fields that together are stable: seller name, listing title, price, and category, hashed into one column.

Put a unique constraint on that key in the database, not just a check in application code. Application-level dedup checks race under concurrency exactly the way my 2am incident did. A database-level unique constraint doesn’t.

Upsert, don’t insert

Once you have a natural key, switch the write path from INSERT to an upsert: INSERT ... ON CONFLICT (natural_key) DO UPDATE in Postgres, MERGE in most other engines, or a manual read-then-write in systems that don’t support atomic upserts natively.

This does two things. It makes reruns safe (the second run just updates the existing row instead of adding a new one), and it gives you free change tracking, since the DO UPDATE clause runs on every rerun even when nothing changed. If you want to know when a price actually changed versus when the scraper simply ran again, you need one more piece.

Content hashing catches the sneaky duplicates

Natural keys solve exact duplicates. They don’t solve the case where the source system doesn’t give you a stable identifier at all, which happens more than you’d expect with scraped content: forum posts, comment threads, classified ads without persistent IDs, PDF-derived text.

For those, hash the normalized content (strip whitespace variance, lowercase, drop volatile fields like “viewed 12 times” counters) and store the hash alongside the row. On the next run, compute the hash again before writing. If it matches an existing hash for that source URL, skip the write entirely rather than upserting an identical row under a new synthetic key.

This also solves a second, quieter problem: sites that serve slightly different HTML on every request (randomized ad slots, A/B test markup, session-specific tracking pixels) which would otherwise make every scrape look like new content even when nothing meaningful changed. Normalize before you hash, or you’ll dedup nothing.

Checkpoints and watermarks, not “just rerun from scratch”

For paginated or incremental scrapes, idempotency also applies to where the job starts, not just how it writes. A scraper that always starts from page 1 and walks forward is safe to rerun by design, if slow. A scraper that resumes from “the last page we saw” needs that watermark stored somewhere durable, and it needs to be updated only after a page’s rows are confirmed written, not before.

I’ve seen the watermark get updated at the start of processing a page, as a kind of progress marker, with the actual database write happening after. If the job dies between those two steps, the watermark says “done” but the row never landed. On rerun, that page gets silently skipped rather than duplicated, which is arguably worse: it’s data loss that produces no error and no duplicate to notice. Update the watermark after the write commits, in the same transaction if your storage layer allows it.

The orchestration layer will retry you whether you like it or not

Airflow, cron with a wrapper, Kubernetes CronJobs, whatever you’re running on, the orchestration layer has its own opinions about retries that are independent of your scraper’s logic. Timeouts, missed heartbeats, and pod evictions all trigger reruns you didn’t ask for. Treating “the scraper only runs once per schedule” as an assumption is how the 2am incident happened.

The practical fix is a job lock: before a run starts, it checks for an active lock keyed to that job and time window (a row in a job_runs table with a status column works fine, you don’t need anything fancier). If a lock already exists and isn’t stale, the new invocation exits immediately instead of scraping in parallel with itself. Set a lock timeout longer than your slowest expected run, or you’ll block legitimate reruns after real failures.

Idempotency keys for the write path

For scrapers that fan out into a queue (scrape workers push extracted records onto Kafka, SQS, or similar, and a separate consumer writes to the database), the natural-key upsert has to happen at the consumer, not just be assumed from the producer side. Queues generally guarantee at-least-once delivery, not exactly-once, which means the same message can arrive at the consumer twice on its own, with no scraper rerun involved at all.

Attach an idempotency key to each message (the natural key plus a run ID works well) and have the consumer check it against a short-lived dedup store, or rely on the same database-level unique constraint from the upsert step. The queue’s delivery guarantee and your database’s write guarantee are two separate problems. Solving one doesn’t solve the other.

Test it by actually rerunning it

None of this is verified by reading the code. The only real test is running the job twice against the same source data and diffing row counts before and after. Run it once, note the count. Run it again immediately, with no changes on the source side. If the count moved, something in the write path isn’t idempotent yet, and it’s worth finding which layer failed: no unique constraint, no content hash, a race in the lock, or a watermark updated too early.

Do this in a staging environment against a snapshot of real target pages, not synthetic test fixtures. Fixtures tend to be too clean to expose the messy cases, like a source page that renders with a different ad slot ID on every load, that cause false “new content” detections in production.

A short checklist

  • Unique constraint on a natural key or content hash, enforced by the database, not just application code
  • Writes go through upsert, not raw insert
  • Watermarks and checkpoints update only after the write commits
  • A job lock keyed to the schedule window, with a timeout longer than the slowest normal run
  • Consumer-side dedup if the pipeline uses a queue between scraping and storage
  • A rerun test, done for real against live or snapshotted source pages, not assumed from reading the code

Idempotent scraper design is not exciting work. It doesn’t show up in a demo and nobody notices it when it’s working, which is the point. The scrapers that end up needing a 2am incident review are the ones where this got skipped because the pipeline looked fine on every individual run.

We write more on pipeline reliability, proxy infrastructure, and the operational side of scraping at scale over at Data Research Tools .

Get new guides and videos first — join the Telegram channel.

Comments

Leave a Reply

Your email address will not be published. Required fields are marked *