How to choose a data source before you write a scraper

Every scraper I have thrown away and rebuilt, I threw away because of the source I picked. Never once because the parsing code was bad. Parsing is a day of work. The source decides your infrastructure bill, your failure rate, and how many mornings you spend repairing something that worked fine yesterday.

Most people make that decision in about four seconds. They open the page, see data on it, and start writing selectors.

Spend ten minutes on it instead.

The order to check

Four surfaces on a typical site, ranked. Stop at the first one carrying the fields you actually need.

  1. A documented api.
  2. An undocumented json endpoint the page calls for itself.
  3. A bulk export, a feed, or a sitemap.
  4. The rendered html.

That ranking is by expected lifespan, and it ranks them by cost per record too. The surfaces with a contract behind them ship the least junk down the wire.

Five minutes on the api question

People skip this because they assume an api will be paid or gated. Check anyway. Look for a developer link in the footer, try docs. and developer. on the domain, search the company name with the word api. Plenty of businesses publish one and never surface it anywhere obvious.

When you find one, resist the urge to celebrate. Open the field reference and compare it against the record you need, field by field, before you write a line. An api missing the two fields your product depends on burns a week before you notice, usually after the ingestion layer is already built around the wrong shape.

Read the quota with the same suspicion. A limit of one request per second caps you at about 86,000 records a day, whatever hardware you throw at it.

The endpoint the page calls for itself

No public api, so look at what the page is already doing on your behalf. Open devtools, network tab, filter to fetch and xhr, reload. On anything built in the last several years you will watch the page fetch its own data as json and assemble the markup from those responses in the browser.

That json is the better surface for one plain reason: it has a schema. Somebody on the other side treats it as the interface between their front end and their back end, so it moves on their release cycle. Layout moves on their design cycle, and design ships far more often than schemas do.

Things to check while you are in there. Whether the response needs a header the page sets or a session cookie. How it paginates, normally an offset or a cursor. Whether the payload carries more fields than the page renders, which it very often does.

This technique changes none of the rules, though. Same server, same terms of service, same rate limits, same defenses in front of it. Finding the endpoint does nothing to make your traffic invisible. It lets you ask for the same data more cheaply, which the operator generally prefers, since one json response costs them a fraction of a full page render. Read the terms and respect the published limits.

Files beat pages

Before you fetch anything page by page, check whether the dataset already exists as a file. Government portals, regulators and a lot of research bodies publish flat dumps: csv, json lines, sometimes an entire database. Failing that, an xml sitemap hands you the url list without crawling for it, and an rss feed hands you new items without polling. Neither of those gives you the data on the page, but both delete the discovery half of the job, which is frequently the larger half of the bill.

The mobile template nobody checks

This is the one I think is most underrated. Try the mobile version.

Plenty of sites still keep a separate lighter template, sometimes on an m. subdomain, sometimes served on the same url to a phone user agent. It exists because somebody wanted the page usable on a bad connection, so it was built small. Small usually means server rendered, with a payload that can run a tenth the weight of the desktop page.

If desktop needs a headless browser and mobile does not, that single difference decides the size of the machine you rent.

Two caveats. Mobile templates sometimes carry fewer fields, so check parity first. And separate mobile sites are being retired in favour of responsive layouts, so the trick is fading.

The ten minute test

The question that sets your entire tooling budget is whether the page is assembled on the server or in the browser. You can answer it without writing code.

Pick a string you can see on the page, something unique. A price, a product name, a review count. Fetch the raw page with curl and search the response for that string.

If it is there, the page is server rendered and an http client is the whole answer.

If it is missing, look for a script tag holding a large json blob. Frameworks routinely embed the full page state that way, and if your string is in there, parse the blob rather than the markup.

If it is in neither place, the page is assembled client side, so go back to the network tab and find the call that fetched it. That call is your real source. The browser was only ever the thing making the call for you.

Rendering the page yourself is the last resort. It should feel like a defeat.

What that answer costs

An http client pulling json holds a few megabytes of memory per worker, so hundreds fit on one modest box. A headless browser instance wants between 300 MB and a gigabyte of ram, plus real cpu for layout and script execution. The machine running four hundred http workers runs maybe eight browsers.

Bandwidth is where it hurts if you buy proxy traffic by the gigabyte. A json request pulls a few kilobytes. Loading the same page in a browser pulls html, css, javascript bundles, fonts, images, and every tracking pixel anybody ever added. Two orders of magnitude more traffic for the same fields is normal.

Html first is a subscription you never cancel

Go straight to selectors and you have signed up for maintenance with no end date, arriving on somebody else’s release schedule.

Layout changes break you. That one is obvious. Less obvious: many modern css class names are generated at build time, so they rotate on every deploy whether or not anything visible moved. A selector written against one of those is a selector written against a build hash.

Split tests break you in the nastiest way, because they only break part of your traffic. Half your requests parse, half return empty, and the symptom looks like a proxy problem. I watched a team chase that for three days.

The tax is per target, permanently. Ten sites parsed from html is ten things failing independently on ten calendars you do not control.

Terms and limits are part of the choice

Two things to read before you build, both of which take longer to fix than to check.

The terms of service, specifically what they say about automated access and reuse of the data. Some sites permit crawling and forbid redistribution. Some forbid both. Some run a paid api precisely because they want automated users on a metered channel, and the honest answer there is to pay for it. If the terms say no, the answer is no, and technical ease has no bearing on it.

Then the published limits. robots.txt tells you which paths the site asks crawlers to stay out of, and it sometimes carries a crawl delay. Ignoring a stated preference is how a source that worked for a year suddenly stops working for everybody.

Check what kind of data it is, too. Personal data sits in a different legal category from product prices in most jurisdictions, and public visibility does not change the category.

The source I picked badly

I needed pricing across about a dozen vendors, and there was an aggregator carrying all of them. One site, one layout, one pagination scheme. One scraper instead of twelve. I took it without thinking very hard.

It worked, which was the problem.

Staleness showed up first, months in. The aggregator refreshed some vendors daily and others whenever it felt like it, and nothing on the page said which. Records three weeks old sat beside records from that morning wearing identical timestamps, because the timestamp was mine rather than theirs.

Then the normalisation. They rounded prices, mapped vendor plan names onto their own categories, and quietly dropped anything that did not fit their schema. What I had was one company’s opinion of the market with the edges trimmed off, and I had been calling it the market.

The thing that killed it was an upstream change on their side. A whole segment of my numbers shifted overnight with no change in my code and nothing in my logs. A customer facing dashboard had been wrong for a stretch I could not pin down afterwards.

The rebuild was eleven separate scrapers against the vendors themselves, four of which had public apis I never bothered to look for. Six weeks of stored data went in the bin because I could not tell which rows were real.

One convenient aggregator puts a company you have no relationship with between you and the truth, and you find out when their editorial decisions leak into your numbers.

Where this rule breaks

The honest caveat. I said json endpoints are more stable than html, and that is true on average and nowhere near a guarantee. I have had an internal endpoint change its response shape twice in a month on a site whose visible layout had not moved in a year. Internal apis have no external users and therefore no compatibility obligation, so nobody deprecates gracefully when the only client is their own front end.

That makes it a bet, and you cover a bet with monitoring. Validate the shape of every response against what you expect, alert when a field disappears or a type changes, and you hear about it within the hour rather than from a customer next month.

More on proxy infrastructure, pipeline design and the tools I actually run is at Data Research Tools.

Get new guides and videos first — join the Telegram channel.

Comments

Leave a Reply

Your email address will not be published. Required fields are marked *