Your cart is currently empty!
Building a scraper someone else can actually run
Most scrapers die the same way. Not from a ban, not from a lawsuit, but from the person who wrote it leaving the company, and nobody else being able to figure out why it stopped working three weeks later. We run scraping pipelines across multiple targets and proxy pools day to day, and the single biggest cost we’ve seen isn’t infrastructure. It’s the hours spent reverse-engineering a script that only made sense to the person who wrote it, at the time they wrote it.
A scraper that runs once on your laptop is a script. A scraper that keeps running after you’ve moved on to something else is infrastructure. The gap between the two is not clever code. It’s a handful of boring decisions made early, about structure, logging, and failure handling, that determine whether the next person can open the repo and actually do something with it.
Separate what changes from what doesn’t
Every scraper has two kinds of logic mixed together: the part that’s specific to one site (which selector holds the price, which endpoint returns the listing) and the part that’s the same no matter what you’re scraping (retry behavior, proxy assignment, output format, logging).
When these live in the same function, every site change forces someone to read code they don’t own. The fix is a config layer. Site-specific values, selectors, URL patterns, pagination rules, field mappings, should live in a config file or database row, not inline in the scraping logic. The scraper itself becomes a generic engine that takes a config and executes it.
This matters more than it sounds like it should, because the person maintaining this six months from now is almost never the person who understands the target site’s HTML. They understand Python, or the pipeline, or the proxy layer. If fixing a broken selector means editing a YAML file instead of tracing through a class hierarchy, that person can fix it in five minutes instead of opening a ticket.
Assume selectors will break, because they will
Sites change their markup. Class names get regenerated by build tools, layouts get A/B tested, entire page structures get rewritten for a redesign. A scraper that treats a CSS selector as a permanent contract is a scraper that fails silently the day someone on the target site ships a frontend update.
The practical response is to make failure visible and specific. If a selector returns nothing, that should raise a distinct, loggable error, not get treated as “field is empty” and passed downstream as null. We’ve seen pipelines that quietly filled a database with empty price fields for two weeks because the scraper treated a missing selector match the same as a legitimately empty value. Nobody noticed until a downstream report looked wrong.
Where you can, prefer selectors that target structural or semantic attributes over ones that target generated class names, since those tend to survive redesigns longer. But don’t over-engineer this. No selector strategy survives a full site rewrite. What you’re really building is fast detection when it breaks, not a selector that never breaks.
Build the retry and backoff logic once, correctly
Every scraper needs retry logic for network failures, rate limit responses, and transient errors. The mistake we see constantly is retry logic copy-pasted into every scraper, slightly differently each time, with no shared understanding of what counts as retryable versus what counts as a real failure that should stop the run.
Write this once as a shared component: exponential backoff with jitter, a cap on retry count, and a clear distinction between “the network hiccuped, try again” and “the target is returning something that means stop, like a persistent block page or a change in response structure that suggests the whole approach needs revisiting.” A scraper that keeps retrying against a page that’s telling it to stop isn’t resilient, it’s just hammering a server that has already told it no. That’s bad practice regardless of whether it eventually gets detected, and it’s the kind of behavior that gets an IP range or account flagged for everyone using it, not just you.
The retry component should also record why it retried. A pile of retry logs with no reason attached is close to useless when you’re trying to figure out whether a target changed something or your proxy pool degraded.
Logging is for the person who isn’t you
The most common failure mode in inherited scrapers isn’t bad logic, it’s logic with no visibility into what it’s doing. A scraper that fails and prints nothing but a stack trace forces the next person to reproduce the failure locally before they can even start debugging it.
Log at the level of decisions, not just errors: which config was loaded, which proxy or session was used for a given request, what the response status was, how many items were parsed successfully versus skipped, and why anything was skipped. This doesn’t need to be elaborate. A structured log line per request, written to something searchable, turns a two-hour debugging session into a five-minute grep.
Separate operational logs from data output. The scraper’s job is to produce clean records; its logs are a different stream entirely, meant for humans, not for the pipeline downstream. Mixing the two makes both harder to use.
Proxies and sessions are infrastructure, not an afterthought
If a scraper is going to run repeatedly against the same target, how it manages IPs and sessions needs to be a first-class part of the design, not a requests.get() call with a random proxy bolted on. That means deciding upfront how sessions are assigned to identities, how rotation is triggered, and what happens when a proxy in the pool is dead or rate limited by the target.
None of this makes a scraper undetectable, and no configuration of proxies or headers guarantees a site won’t recognize automated traffic. What good proxy management actually buys you is stability: fewer failed requests from exhausted or flagged IPs, and a pipeline that degrades gracefully (skip and log, don’t crash the whole run) when one part of the pool stops working. If you’re evaluating proxy providers for this kind of work, judge them on measurable things, connection stability, documented pool size, response latency under load, not marketing claims about being undetectable, because that claim isn’t something any vendor can actually back up.
Respect published rate limits and robots.txt where they exist. This isn’t just a compliance checkbox, it’s also what keeps a shared proxy pool usable, since aggressive request rates from one script degrade the pool for every other job running on it.
Documentation that isn’t a wall of text
A maintainable scraper needs a short README that answers four questions: what does this scrape, how do you run it, where does the config live, and what does a healthy run look like versus a broken one. That last one is the part people skip, and it’s the one that saves the most time. If the next person doesn’t know that “200 items in 40 minutes” is normal and “12 items in 40 minutes” means something’s wrong, they’ll waste hours figuring that out from scratch, or worse, ship broken data because it looked plausible.
Skip long docstrings describing what each function does structurally. Code with reasonable names already communicates that. Spend the documentation budget on the things that aren’t visible from reading the code: why a particular wait time was chosen, why one field is parsed from a different endpoint than the rest, what happened the last time this broke.
Testing without depending on the live site
You can’t unit test against a live target reliably, since the site changes and you don’t want tests hitting it constantly anyway. Save representative HTML or JSON responses as fixtures and test your parsing logic against those. This catches most regressions, parsing errors from a code change, a field that got dropped, without needing network access, and it gives you a fast way to confirm a fix actually works before deploying it against the real target.
The actual measure of maintainability
None of this is exciting. Config separation, structured logging, shared retry logic, and a short README aren’t the parts of a scraping project anyone shows off. But they’re the difference between a pipeline that survives turnover and one that gets rewritten from scratch every time someone new inherits it. If you’re building scraping infrastructure that’s meant to outlast the person who wrote it, that’s the actual bar, not clever selectors, but whether someone else can open the repo, understand what broke, and fix it without you.
If you’re building out scraping infrastructure and want more on framework choices, proxy setup, and pipeline design from people running it in production, check out the rest of what we’ve written at Data Research Tools.
Get new guides and videos first — join the Telegram channel.
Leave a Reply