Your cart is currently empty!
Validating scraped records before they hit your database
The gap between “it scraped” and “it’s correct”
A scraper that returns 200 OK and a parsed object is not the same thing as a scraper that returned a correct record. Anyone who has run a pipeline past a few million rows has seen the failure mode: the job finishes green, the row count looks right, and three weeks later someone finds that 8% of the “price” field is actually a shipping estimate because a site changed a CSS class and the parser quietly grabbed the wrong <span>.
The scraper didn’t crash. It didn’t even throw a warning. It just picked up the wrong node and wrote a plausible-looking number into a plausible-looking column. That’s the core problem with scraped data: most failures aren’t exceptions, they’re silent substitutions. A validation layer exists to catch exactly that class of error, before it becomes a row in production.
Why this has to live between extraction and storage
Teams often bolt validation onto the end of the pipeline, as a query that runs against the database after the fact. That’s too late. Once a bad batch lands in the same table as good data, you’ve mixed the two, and cleaning it up means writing a reconciliation script that has to guess which rows came from the bad run. If validation instead sits between the parser and the insert, a bad batch never reaches the table. It gets logged, held, or routed to a review queue, and the production data stays trustworthy.
This means validation is a pipeline stage, not a database constraint you add later. Constraints (NOT NULL, foreign keys, unique indexes) are useful as a last line of defense, but they only catch structural violations. They will not catch a price field that’s technically a valid decimal but is off by two orders of magnitude because the parser read the wrong element.
Schema validation: catching shape before catching meaning
The first layer is structural. Every record coming out of a parser should be checked against a schema before it’s queued for insert: field presence, type, and basic bounds. If you’re on Python, Pydantic or jsonschema do this cheaply. The point isn’t the library, it’s the discipline of defining what a valid record looks like as code, not as a mental model you re-derive every time something breaks.
Concretely, this means:
- Required fields are actually present, not just present as empty strings or
nullmasquerading as valid data - Numeric fields parse as numbers, not as strings that happen to look numeric
- Enums (currency codes, status fields, category labels) are checked against a known set, not accepted as free text
- String fields have sane length bounds, so a scraper that accidentally captured an entire page body into a
titlefield gets flagged instead of silently truncated by a database column limit
Schema validation is cheap and catches maybe 60% of real-world scraping failures: missing fields when a page returns a partial render, type mismatches when a site A/B tests a different markup structure, truncated strings when a request gets cut off mid-response. It will not catch the harder failures, where the record is well-formed but wrong.
Semantic validation: catching wrong when it looks right
This is the layer that actually matters for data quality, and it’s the one most pipelines skip because it’s harder to write generically. Semantic checks encode what you know about the domain, not what a schema library can infer from types.
A few patterns that hold across most scraping targets:
Range and plausibility checks. If you’re scraping product prices in SGD, a value of 4 or 400,000 for a consumer product is a red flag worth a second look, not an automatic reject, but a flag. Set soft bounds per category if you have them, and log outliers to a review table instead of silently accepting or silently dropping them.
Cross-field consistency. If a record has a “discount price” higher than its “list price,” something upstream broke, most likely the parser grabbed two different DOM nodes that used to be in a fixed order and no longer are. If a record has a “posted date” that’s in the future relative to your scrape time, your date parser is misreading a format (this happens constantly with sites that mix DD/MM and MM/DD without a clear locale signal).
Referential plausibility. If you’re scraping listings that should map to a fixed set of categories, sellers, or locations, check new values against your known set before accepting them as new entities. A sudden new “category” that’s actually a JavaScript error string rendered into the DOM is a common artifact of scraping a page mid-render.
Duplicate and near-duplicate detection. Re-scraping the same listing across runs is normal. What you want to catch is the same underlying entity showing up with a different ID because the site rotates listing IDs, or the same content getting inserted twice because a retry succeeded on a request that had already succeeded once. A content hash (normalize whitespace, hash key fields) checked against recent inserts before writing catches this at negligible cost.
Handling drift, not just errors
Sites change their markup. This is not an edge case, it’s the normal operating condition of any scraper that runs for more than a few weeks. The validation layer’s job here is not just to catch bad records, it’s to catch the moment drift starts, so you find out from a monitoring alert rather than from a customer complaint.
The practical version of this: track the pass rate of your validation layer per source, per run, as a metric over time. If a source normally validates at 98% and a run comes in at 61%, that’s not 37% of records being individually wrong, it’s almost always one selector breaking across the whole batch. A sudden drop in validation pass rate is a far more useful signal than any single failed record, because it tells you where to look in the parser instead of making you audit rows one at a time.
Set a threshold (we’ve found somewhere around a 10-15 point drop in pass rate is a reasonable alert trigger, though this depends heavily on how noisy your source normally is) and treat crossing it as “stop and check the parser,” not “let the run finish and clean up later.”
Where bad records should go, not just how to catch them
A validation failure is information, and throwing it away loses that information. Records that fail validation should land in a dead-letter table or queue with the raw payload, the specific check that failed, and a timestamp, not just get dropped from the pipeline silently. This gives you three things: a way to reprocess records once you fix the parser, a way to measure how often each failure mode actually occurs, and an audit trail if someone downstream asks why a number looks off.
This is also where you build the feedback loop that improves the parser over time. If “cross-field price check” is triggering on 5% of records from one source, that’s not necessarily 5% bad scrapes, it might be a legitimate promotional pricing pattern your rule doesn’t account for yet. Reviewing the dead-letter queue periodically is how validation rules stay accurate instead of becoming a source of false rejects that quietly shrink your dataset.
Keeping this proportional to the pipeline
None of this requires a heavyweight data quality platform for most pipelines. A schema check, a handful of domain-specific semantic rules, a content hash for dedup, and a dead-letter table cover the overwhelming majority of real failures we’ve seen in production scraping. The mistake isn’t skipping some exotic validation technique, it’s skipping validation entirely and finding out about a broken selector from a downstream report instead of from your own pipeline.
If you’re building or auditing a scraping pipeline and want to see how the rest of the infrastructure, orchestration, proxy layers, and framework choices fit around this, we cover it in more depth on the site.
Explore more scraping infrastructure guides and tool breakdowns on the Data Research Tools home page.
Get new guides and videos first — join the Telegram channel.
Leave a Reply