Your cart is currently empty!
Running forty scrapers without forty codebases
If you’ve ever inherited a scraping operation, you’ve probably seen the pattern: a folder full of scripts, each one named after a site, each one copy-pasted from the last one and then edited until it worked. Forty targets means forty scripts, forty sets of retry logic, forty places where a proxy pool got hardcoded, and forty places where nobody remembers why a particular sleep() call is there. It runs, until it doesn’t, and then nobody wants to touch it.
This is not a scale problem you hit at some huge number. It shows up around scraper five or six, when you notice you’re fixing the same bug in three different files because you fixed it in the first one and forgot to port the fix. By scraper twenty it’s not a nuisance, it’s the reason your team spends more time on maintenance than on new coverage.
The fix isn’t a framework you buy. It’s a decision about what belongs in code and what belongs in configuration, made early and enforced consistently.
What actually varies between targets
Look closely at forty scrapers pulling data from forty different sites and most of what’s “different” is really the same handful of concerns, expressed differently:
- Where the target data lives (a URL pattern, a sitemap, a paginated listing)
- How the page is structured (which selectors or JSON paths hold the fields you want)
- What counts as a rate limit or a block (a status code, a redirect to a challenge page, a specific response body)
- How often the site should be hit and from where
None of that is really “logic.” It’s data about a target. The actual logic, fetch a URL, retry on failure, parse a response, normalize a record, write it to storage, is identical across all forty. The mistake most teams make is letting the parts that should be data (a CSS selector, a rate limit, a base URL) get buried inside the parts that should be code (a Python file with control flow).
A shared core, not a shared script
The way out is a single core library that does the boring, repeatable 80%: HTTP fetching with retry and backoff, proxy assignment, response caching, structured logging, and a normalized output schema. Every scraper imports this core. None of them reimplement it.
What’s left per target is a small definition, not a program. In practice this is often a single config file or a thin adapter class with three or four methods: how to build the request, how to extract fields from a response, and how to detect that a response wasn’t actually valid data (a soft block, an empty result set, a login wall). That adapter might be forty or fifty lines. It is not a scraper in the sense of “a standalone thing that runs.” It’s a description that the core engine consumes.
This is the difference between forty codebases and one codebase with forty configs. When you fix a retry bug in the core, you fix it for all forty targets at once. When a proxy provider changes its authentication format, you change it in one place. The blast radius of any given change becomes predictable, which is the entire point.
Where config-driven scraping breaks down
Config-driven design works well right up until a target does something that genuinely needs procedural logic: a multi-step login flow, a page that renders fields conditionally based on earlier responses, a site where pagination depends on values returned in a previous page rather than a predictable URL pattern. Trying to express that as a static config file usually produces something uglier than just writing code, a config format that’s secretly a badly-designed programming language.
The honest answer here is to let the adapter be code when the target needs code, but keep it isolated. A target that needs a five-step interaction still shouldn’t touch retry logic, proxy selection, or storage. It writes its own extraction and flow logic and still calls into the shared core for everything else. You end up with a few “heavy” adapters among a majority of simple ones, which is a much smaller maintenance surface than forty independent scripts of unknown quality.
Proxy and fingerprint handling belongs in infrastructure, not in scripts
One of the fastest ways to end up with forty divergent codebases is letting proxy logic live inside each scraper. One script rotates proxies every request, another rotates every ten minutes, a third has a hardcoded IP that someone set for a one-off debug session and never removed. None of this is a scraping problem, it’s a networking problem, and it should be solved once, below the scraper layer.
In practice that means the core library asks a proxy manager for an egress point and gets back a working connection, without the scraper code knowing or caring whether that’s a residential IP, a datacenter IP, or a mobile carrier IP. Session stickiness (keeping the same IP for a sequence of related requests), rotation intervals, and per-target IP pools are configuration the proxy layer enforces, not something each scraper decides for itself.
The same goes for browser fingerprinting when a target requires a real browser context rather than a plain HTTP client. Fingerprint consistency (matching a real device profile across headers, TLS handshake, and JS-visible browser properties) is something anti-detect browser tooling and proxy providers build entire products around, and it’s worth understanding briefly why it matters: modern bot detection doesn’t just look at the IP, it correlates network-level signals with browser-level signals and flags mismatches between them. That’s infrastructure work, done by dedicated tooling and reviewed periodically, not something to reinvent per scraper. No proxy or browser fingerprinting setup makes a scraper undetectable, and treating it that way is how operations end up surprised when a target changes its detection logic.
Monitoring forty things instead of one
A second, quieter cost of the copy-paste approach is that nobody actually knows the state of forty independent scripts. Each one logs differently, if it logs at all. Failures show up as “the CSV is empty this morning” rather than an alert.
With a shared core, monitoring becomes a byproduct rather than a separate project. Every run reports through the same structure: rows extracted, request count, error rate, and whatever the target-specific adapter flagged as a soft failure. A dashboard that shows forty rows, one per target, with a last-successful-run timestamp and an error rate, tells you in ten seconds which targets are actually healthy. That’s not possible when every scraper is a bespoke script with bespoke logging, or none at all.
Respecting the target, at scale
Running many scrapers against many sites is also where it becomes easy to lose track of what’s reasonable. A rate limit that made sense for one script, hand-tuned by someone watching logs, doesn’t automatically transfer when it’s one of forty jobs sharing a proxy pool and a scheduler. Centralizing request pacing per target, rather than per script, is also how you avoid accidentally hammering a site because three different scrapers independently decided to hit it. It’s worth treating a target’s published terms of service and robots.txt as an input to that per-target configuration, not an afterthought, and steering clear of anything gated behind a login or paywall that wasn’t meant to be machine-read.
When a new codebase is actually the right call
None of this means one codebase forever. If a new class of target is structurally different enough, JavaScript-heavy single page apps versus static HTML, or a GraphQL API versus scraped markup, it can be worth a second core built for that shape of problem, with its own adapters. The failure mode to avoid isn’t “more than one codebase.” It’s a codebase per target, which is what happens by default when nobody makes the shared-core decision on purpose.
The real test of whether your scraping operation is a system or a pile of scripts is simple: when a target’s HTML structure changes, how many files do you have to touch, and how long does it take you to know it broke in the first place. If the answer is “one file” and “a dashboard told me this morning,” you’ve built infrastructure. If it’s “grep for the site name across the repo” and “a customer complained,” you’re still in script territory, no matter how many scrapers you’re running.
If you’re building out proxy infrastructure, orchestration, or fingerprinting defenses to support an operation like this, you can find more breakdowns like this one on the Data Research Tools home page.
Get new guides and videos first — join the Telegram channel.
Leave a Reply