Your cart is currently empty!
Detecting when a scrape silently returned nothing
A scraper that throws an exception is a good day. Someone sees the stack trace, someone fixes it, the pipeline moves on. The bad day is the one where the job finishes clean, logs say success, HTTP status is 200, and the output file has three rows in it instead of three thousand. Nothing crashed. Nothing paged anyone. The data just quietly wasn’t there.
This is the failure mode that does the most damage in production scraping, because it doesn’t announce itself. A crash is loud. An empty result set dressed up as a successful run is silent, and silent failures compound: the next run overwrites yesterday’s good data with today’s empty set, a dashboard downstream shows a gap that gets read as “no news today” instead of “the pipeline broke,” and by the time a human notices, a week of data is missing and nobody can say exactly when it started.
Why “it ran” isn’t the same as “it worked”
Most scraping frameworks treat success as “the request completed and the code didn’t throw.” That’s a network-layer definition of success. It says nothing about whether the page you got back is the page you wanted.
There are several ways a request can complete cleanly and still hand you nothing useful:
A challenge page instead of content. Many anti-bot systems don’t return a 403 or a 429 when they’re suspicious of a request. They return 200, with a page that looks like HTML but contains a JavaScript challenge, a CAPTCHA widget, or an “checking your browser” interstitial. If your parser is looking for a specific CSS selector and that selector isn’t on the challenge page, you get an empty result, not an error. The status code lied by omission.
Selector drift. A site redesigns a product listing page. The container div that used to hold .item-card now holds .product-tile. The request succeeds, the HTML is real content, and your parser still returns zero rows because it’s looking for a class name that no longer exists. This is the most common cause of silent-empty in practice, and it has nothing to do with blocking at all.
JavaScript that never rendered. If a page builds its content client-side and your fetcher is a plain HTTP client rather than a headless browser, you get the empty shell HTML and none of the injected content. Status 200, body present, target data absent.
Soft rate limiting. Some sites respond to excessive request volume not with a block but with a stub response, a cached placeholder, or a page that omits the dynamic sections entirely while keeping the static shell intact. It’s designed to look normal to a script that isn’t checking closely.
All four of these produce the same observable symptom from a pipeline’s point of view: a request that “succeeded” and a parse that returned nothing. The fix isn’t a smarter scraper. It’s a pipeline that stops trusting HTTP status as a proxy for data quality.
Check the shape before you check the content
The cheapest and highest-value guard is a shape check that runs immediately after parsing and before anything touches storage. It doesn’t need to understand the data, it just needs to know what “normal” looks like in aggregate.
Row count against a floor. If a job has returned somewhere between 800 and 1,200 records on every run for the past month, a run that returns 4 records should never reach the database unflagged, even if every one of those 4 records is individually valid.
Byte size of the raw response, before parsing. A challenge page and a real listing page are often wildly different in size. Logging response size per request and comparing it against a rolling baseline catches soft blocks even when the parser itself doesn’t error out.
Presence of expected markers. Rather than trusting that a selector match means the page is real, check for a handful of strings or elements that should exist on a genuine page but wouldn’t exist on a challenge page, and vice versa: strings that only show up on interstitials (captcha, verify you are human, checking your browser) and should never appear in a healthy scrape.
Schema completeness. If your target record has a price, a title, and a URL, a “successful” parse that returns records with all three fields null isn’t successful. Validate required fields exist and are non-empty before the record is written, not after.
None of these checks need to be clever. They need to run every time, on every batch, before data lands anywhere downstream systems will read from.
Baselines matter more than absolutes
A hard rule like “reject if zero rows” catches the extreme case but misses the more common one: a job that used to return 1,000 rows and now returns 40. Forty is not zero. It will pass a naive check and still represent a pipeline that’s mostly broken.
The more durable approach is comparing each run against a rolling baseline built from recent history, not a fixed number set once and forgotten. If a job’s typical output sits in a known range, a run that falls outside a few standard deviations from that range is worth flagging even if it’s technically nonzero. Sites change layout gradually sometimes, and a slow bleed from 1,000 rows to 600 over two weeks is exactly the kind of thing a static threshold will never catch but a trend comparison will.
This also has to account for legitimate variance. A site that only lists 200 items on a Sunday isn’t broken, it’s just Sunday. Baselines that ignore day-of-week or seasonal patterns generate false alarms, and false alarms are how teams end up ignoring the alert channel entirely.
Where the check lives in the pipeline
Putting the validation step after storage means bad data has already landed before anyone finds out. The check needs to sit between parse and write: parse the response, run the shape and baseline checks, and only then commit to the database or the file store. A run that fails validation should be quarantined, not silently dropped and not silently accepted. Quarantine means the raw response and the parse output both get saved somewhere inspectable, so a human or a secondary process can look at exactly what came back and decide whether it’s a selector problem, a blocking problem, or a real change in the underlying data.
Retrying immediately on failure is tempting but often wastes the run. If a request got a challenge page because of how it looked (missing headers, a proxy IP with a poor reputation, a fingerprint that doesn’t match a real browser), retrying the identical request from the identical source usually produces the identical result. A better pattern is to flag the failure, stop that job from overwriting good historical data, and let a separate diagnostic pass investigate root cause rather than burning request volume in a retry loop that can’t succeed.
Alerting without training people to ignore it
An alert that fires on every run with slightly lower volume than yesterday will get muted within a week. The signal worth paging someone for is a sustained pattern: three consecutive runs below baseline, not one. A single anomalous run is worth logging. A trend is worth waking someone up for.
It also helps to separate “this looks like blocking” from “this looks like a broken selector” in the alert itself, because the fix is completely different. A blocking signature usually shows response size and structure changing sharply while request success rate (in the HTTP sense) stays high. A selector break usually shows response size staying normal while the specific fields being extracted go to zero or null. Logging both signals separately, rather than collapsing everything into one “job failed” bucket, saves the person on call from having to re-diagnose from scratch every time.
The underlying point
A scraping pipeline that only checks “did the request return 200” is checking the wrong layer. The network succeeding and the data being real are two different claims, and production systems that treat them as the same thing find out the hard way, usually when someone downstream asks why a report has been flat for two weeks. Detection has to happen at the data layer: shape, volume against baseline, and required-field completeness, checked before anything is written and quarantined loudly when it doesn’t hold up.
If you’re building out pipeline monitoring, proxy infrastructure, or want a closer look at how scraping frameworks handle failure detection in practice, more breakdowns like this live on the Data Research Tools home page.
Get new guides and videos first — join the Telegram channel.
Leave a Reply