Building a parser that survives a site redesign

A scraper I run pulls listings off a supplier catalogue every night. It ran for eight months without a complaint. Then somebody wrapped the price block in one more container to fix the spacing on mobile, and the price column went blank for nineteen days before anyone noticed.

Nothing failed. The job exited zero every night, the row count held steady, and one column had quietly turned into a wall of empty strings.

That is the shape of every parser incident worth talking about. The fetch worked and the pipeline was green. The data was wrong.

What copy selector actually gives you

Almost everyone starts here. Open devtools, right click the element, copy selector, paste. It works immediately, which is what makes it dangerous.

Read what you pasted. It is a path through containers: a wrapper, a grid, the second child, a class name that is a hash of some styling rules. Every step in that chain is a promise the site never made you.

Add a wrapper and it breaks. Insert an advert slot above the target and the child index shifts, so the selector still matches something and hands back the wrong value. That is the worse outcome, because nothing looks broken.

If the site is built with a css-in-js library, those class names are generated from the style rules. Change a padding value, ship it, and the hash changes. No redesign required.

The test I use now: read the selector aloud. If you cannot say which field it grabs without opening the page in front of you, it describes position. Position is the shortest-lived thing on a web page.

A ranking of anchors by how long they last

You pick an anchor once per field and then live with it for years, so the order is worth memorising.

Structured data first. Most listing and commerce sites publish a JSON-LD block, microdata attributes or Open Graph tags for search engines. That markup is generated from the same database that renders the page, and the site’s rich results depend on it staying correct, so it gets maintained by people who are not thinking about you at all.

Watch it for staleness. I have seen a JSON-LD price hold the old value for several hours after the visible page updated, so check it against the rendered page when you wire the field up.

Second, identifiers and data attributes the site relies on for its own logic. An id its JavaScript queries to attach a handler. A data-testid its test suite depends on. Renaming one of those costs the site something. You can usually tell by reading them: data-product-price is a deliberate hook, class="flex gap-4" is decoration.

Third, the text label sitting next to the value. On a spec table, find the row whose header cell reads Warranty and take the cell beside it. This survives a column reorder and a complete restyle, because you are following the word a human reads.

Labels have their own failure mode. They are copy, and copy gets rewritten. Warranty becomes Cover. A locale switch returns half of them in German. Match against a small accepted set instead of one exact string.

Last, position in the tree. Second cell, third paragraph, nth-child. Use it when the page gives you nothing else, and leave a comment saying so, because that is the field that will break first.

The empty string that cost me nineteen days

I should own the specific mistake, because everything above came out of it.

I had a helper wrapping every extraction. Selector matched, return the text. Selector missed, return "". I wrote it that way so the calling code stayed tidy with no null checks scattered through it. I was pleased with it at the time.

Then the container changed and every price came back as "".

My required-field validation checked for None. An empty string is not None, so validation passed. The column was declared TEXT, so Postgres took it without comment. Row counts stayed normal, the job stayed green, and a dashboard downstream showed a full table with one column of blanks that nobody had reason to look at.

I found it nineteen days later, and only because a reconciliation against a second source came out wrong.

Worse, I was not storing raw pages then, so there was nothing to reparse. Every one of those listings had to be fetched again, through metered mobile lines, at a pace slow enough to avoid getting blocked. That took most of a week. A handful had already rotated off the site, and those rows are still empty.

The bug was one line. My helper collapsed the difference between a field that is empty and a field that is missing, and those are two different facts about the world. It raises now.

Extract wide and keep the bytes

Two habits, both cheap at the time you write them.

Pull more fields than the current requirement asks for. You already paid for the fetch, and pulling another eight values out of a DOM you are already holding costs nothing measurable.

Then store the raw capture. Compressed bytes, sitting next to the URL, the fetch timestamp and the parser version that produced the row.

That single decision changes the price of every future break:

  • With the capture, a redesign is a reparse. Fix the extraction, run it back over stored pages, and history repairs itself at the speed of a batch job. No proxies, no rate limits.
  • Without it, the same redesign is a recrawl. Two weeks of pages fetched again through metered bandwidth at a polite pace, and any listing that has rotated off the site since is gone permanently.

You do not need infinite retention. A rolling window of a few weeks covers realistic detection lag, and one saved sample per page type, kept indefinitely, gives you fixtures to test against.

Assertions belong at the field level

Most parsing code fails soft by default. No match, return None or "", record continues down the pipeline looking perfectly normal. That default is the entire problem.

Give the parser two outcomes. Either it produced a record whose shape it recognises, or it refused and named the field that failed.

The assertions themselves are boring, which is fine:

  • price parses as a number and is greater than zero
  • title is present and under 300 characters
  • currency is one of the four you actually deal in
  • a date parses and does not land in the future

Each is a couple of minutes of work. Together they turn a redesign into a first-record failure instead of a slow drip of blanks nobody sees.

Do not kill the whole run over one bad record, though. Quarantine it: write the record to a rejects table with the raw HTML attached, increment a counter, carry on. A normal night sits well under 1% rejects. The morning after a redesign it is 90-something, and the HTML you need to write the new selector against is already sitting there.

Fill rate, and the alert that actually fires

Fill rate is the proportion of rows where a given field came back empty. One query per field, and the most useful number in the pipeline.

The mistake is comparing it against a fixed threshold. Real data has legitimately sparse fields: a discount that only exists during a promotion, a second image that only some listings carry. Set the alarm at 90% full and those fields page you nightly until someone mutes the channel. Set it at 10% and a field sliding from 99% to 60% never trips.

Compare a field against its own history instead. Trailing median over the last 14 days, per field and per source, with an alert on a deviation of more than a few points. The absolute level carries almost no information. The delta carries all of it.

Splitting by source matters more than it sounds. A global average across twenty sites hides one broken site completely, because the other nineteen hold the number up for it.

A strategy list per field, and recording the winner

Instead of one selector, give each important field an ordered list of extraction strategies. JSON-LD, then the data attribute, then the label-anchored lookup, then the positional guess. First one to return a plausible value wins.

The half of this that people skip is recording which strategy won. Store the strategy name on the record.

That column is the early warning. If strategy one carried 99% of rows for a month and this morning strategy three is carrying them, the page changed. The data is still correct, so nothing else in the stack has any idea. You get a week to fix it properly instead of an afternoon to fix it badly.

Skip the column and your fallbacks absorb the damage silently until the last one dies too, and then you get the outage anyway with no idea when the decay started.

It does cost more code per field, so scope it. Fields somebody would ring you about get the full ladder. Everything else gets one strategy and a loud assertion.

What a resilient parser still does not fix

None of this makes a parser permanent. A site can restructure hard enough to take out every rung at once, and sooner or later one will. The goal is a break that is cheap to repair and visible the day it happens.

It also has nothing to say about what you are allowed to collect. Public pages, robots.txt respected, a request rate that does not hurt the site, and an official API or a bulk feed used whenever one exists. A parser you never had to write is the only one that cannot break.

We write up the rest of this work, including the fill rate queries, over at Data Research Tools.

Get new guides and videos first — join the Telegram channel.

Comments

Leave a Reply

Your email address will not be published. Required fields are marked *