Why we reparse stored HTML instead of refetching the page

The two costs of a scrape

Every scrape has two separate costs baked into it: the cost of getting the bytes off the target site, and the cost of turning those bytes into structured data. Most people who are new to scraping treat these as one step, because in a quick script they usually are: request the page, run a parser on the response, write the result, move on. That works fine for a hundred pages. It falls apart at scale, because the fetch is the expensive, risky, rate-limited part, and the parse is the cheap, deterministic, endlessly repeatable part. Once you notice that asymmetry, you stop treating them as one operation.

Fetching costs proxy bandwidth, which on residential or mobile proxy pools is often billed per gigabyte and can be the largest recurring line item in a scraping operation. It costs request budget against a target that’s watching request volume per IP, per session, or per account. It costs time, since a crawl over millions of URLs with polite delays and rotation can run for days. And it carries risk: every request is a chance to trip a rate limit, get served a challenge page, or burn a proxy that was working fine a minute ago. Parsing, by contrast, is just CPU. You can run a parser against the same stored response a hundred times while iterating on a selector, and it costs nothing but compute time.

Splitting fetch from parse

The fix is architectural, not clever: store the raw response the moment you fetch it, before any extraction happens. The fetch stage’s only job is to get the page and write it somewhere durable, unmodified. The parse stage reads from that store, not from the network, and produces the structured output. They can run on completely different schedules. A crawl might run once a day; the parser that turns that day’s HTML into rows might run a dozen times in an afternoon while someone is debugging a broken selector.

This sounds obvious once you say it out loud, but a surprising number of pipelines skip it because storing raw HTML feels wasteful compared to storing the clean JSON you actually wanted. The tradeoff only looks that way until the first time a parser bug ships to production and you realize the only way to fix the corrupted output is to re-crawl a million pages you already had.

What actually goes in the archive

What you store matters as much as the fact that you store it. At minimum, keep the response body, the final URL after redirects, the HTTP status code, the response headers, and a timestamp for when the fetch happened. If the fetch went through a proxy, record which one, or at least which proxy pool and IP class, since that’s often the fastest way to explain why one batch of pages looks different from another.

For JavaScript-heavy pages fetched with a headless browser, store the rendered DOM after execution, not just the initial response. The headless render is the expensive part of that fetch, often several times slower and heavier than a plain HTTP request, so it’s the render output you want to be able to reparse for free, not just the raw markup that arrived before the page’s scripts ran.

It’s also worth storing pages that come back wrong. A response that’s actually a CAPTCHA page, a block page, or a soft-404 dressed up as a 200 is still worth keeping, tagged as such. Knowing that a batch of “successful” fetches were actually challenge pages is useful for tuning a crawl, and you can’t tell that after the fact if you threw the body away because the status code looked fine.

Why parser bugs stop being crawl problems

Once the raw HTML is sitting in storage, a parser bug becomes a code fix and a rerun, not a fresh crawl. Say a site restructures a price block and your extraction starts pulling the wrong number for a subset of pages. Without an archive, fixing that means re-fetching every affected page, which means burning proxy bandwidth and request budget a second time for content you already had. With an archive, you patch the parser, point it at the stored batch, and regenerate the output in minutes.

This matters more than it sounds like on paper, because parser bugs are common in a way network failures aren’t. Sites change markup, add or drop optional fields, switch from server-rendered to client-rendered blocks, or vary their layout by region or A/B test. A selector that worked against your sample set will eventually miss something in the full corpus. Being able to iterate on extraction logic against a fixed, already-paid-for dataset instead of a live target is the difference between a ten-minute fix and a multi-day re-crawl.

Backfilling new fields without a new crawl

The same mechanism pays off when requirements change rather than break. If a project decides it now needs a field nobody thought to extract three months ago, and you kept the raw pages from that period, you can backfill the new field by reparsing the archive instead of explaining to whoever’s paying the proxy bill why you need to re-crawl three months of history. This only works if the field was actually present in the page at fetch time, obviously. Reparsing can’t recover data that was never in the response, and it can’t reflect anything that’s changed on the live site since the page was captured. It’s a way to get more out of a fetch you already paid for, not a substitute for fetching again when you need current data.

The tradeoff: storage isn’t free, and stored pages go stale

None of this is free. Raw HTML is heavier than the structured output you eventually derive from it, sometimes by an order of magnitude, so storing it at scale means real object storage cost and lifecycle decisions. A common pattern is to compress bodies with gzip on write, which HTML compresses well, and to set a retention window: keep raw pages for a defined period, say enough to cover a couple of parser iteration cycles, then prune down to the parsed output once you’re confident the extraction is stable. What you keep forever is the structured data; what you keep temporarily is the raw archive that let you get there safely.

The other real limit is staleness. A stored page is a snapshot of the target site at the moment it was fetched. Reparsing it a week later gives you a better extraction of that same snapshot, not fresh data. If a price, a stock count, or a headline has changed on the live site since you fetched, no amount of reparsing recovers that, only a new fetch does. It’s worth being explicit about this distinction internally, because it’s easy for “we can reparse instead of refetching” to quietly turn into “we don’t need to refetch,” when the two are only interchangeable for fixing how you read old data, not for keeping current with a moving target.

Debugging with the archive

There’s a secondary use for a raw archive that has nothing to do with backfilling: it’s the fastest way to figure out what actually changed when a scraper that worked yesterday breaks today. Instead of guessing at what a target site did differently, you can diff a newly fetched page against the last stored version and see the structural change directly, whether that’s a renamed class, a moved block, or a switch to a different rendering path entirely. That diff is usually a five-minute investigation. Without an archived baseline, the same investigation means re-fetching a known-good page from history, which you may not be able to do at all if the site has moved on since then.

Where this fits in a pipeline

In practice this means a scraping pipeline is really two systems joined by a store: a fetcher that’s optimized for reliability, proxy rotation, and staying under a target’s rate tolerance, and a parser that’s optimized for iteration speed and correctness. The fetcher writes once and moves on. The parser reads as many times as it needs to, against data that’s already paid for. That separation is what makes it possible to fix extraction bugs, add fields, and debug breakage without treating every mistake as a reason to go back and hit the target site again.

If you’re building or auditing a scraping pipeline and want more on how the fetch and parse layers should be structured, proxy infrastructure that supports it, or how bot detection actually works from the defending side, take a look at what we’ve written on it at Data Research Tools.

Get new guides and videos first — join the Telegram channel.

Comments

Leave a Reply

Your email address will not be published. Required fields are marked *