Your cart is currently empty!
When a site serves your scraper different pages
You load the page in a browser and it looks normal. You run the same URL through your scraper and the response is thinner, or it’s a different layout entirely, or the data you need just isn’t there. Nothing crashed. No 403, no CAPTCHA, no error at all. The request “succeeded” and it still gave you the wrong thing.
This is one of the more frustrating failure modes in scraping because it doesn’t look like a failure. Status code 200, valid HTML, page renders fine. It’s just not the page a real visitor sees. Here’s what’s actually happening on the other end, and how to figure out which version you’re getting instead of guessing.
The request isn’t just a URL
A browser hitting a page and a scraper hitting the same URL are sending very different requests even when the URL string is identical. The server sees headers, TLS handshake characteristics, connection behavior, and often a JavaScript execution environment that a bare HTTP client never provides. Any of those can be used to decide what to serve.
The mechanism that matters most here is server-side branching: the backend looks at signals in the request and picks a response template before it ever sends bytes back. This is different from a block, where the server refuses the request outright. A differential response is the server deciding “this looks like X kind of visitor, give them the X version.” Your scraper isn’t rejected, it’s just categorized.
Common signals that feed that decision:
- User-Agent string. The oldest and weakest signal, but still checked. A default
python-requests/2.31.0UA is an instant tell that no browser sent this request. - Header set and order. Real browsers send a specific, consistent set of headers in a specific order (Accept, Accept-Language, Accept-Encoding, Sec-Fetch-* headers, etc). A scraper using a minimal HTTP client often sends a much shorter list, or sends them in an order no browser produces. Some anti-bot vendors fingerprint header ordering specifically because it’s cheap to check and hard for naive clients to get right.
- TLS and HTTP/2 fingerprinting. The TLS ClientHello (cipher suite order, extensions, supported curves) and the HTTP/2 SETTINGS frame differ between Chrome, Firefox, curl, and most scraping libraries. This is checked before a single byte of your HTTP request is parsed, so it can steer the response before content negotiation even happens.
- Cookie and session state. No-cookie requests, or requests missing a cookie that a real page visit would have set on a prior load, often get routed to a stripped-down or cached fallback version.
- JavaScript execution. If a page’s real content is rendered client-side and your scraper only fetches the raw HTML, you’re not seeing a deliberately different page, you’re seeing the pre-render skeleton. This looks identical to a differential response from the outside but the cause and the fix are different.
None of this requires the site to know it’s talking to your specific scraper. Most of it is bucket sorting: does this request look enough like a browser to get the full experience, or does it get the cheap path.
Why sites actually do this
It’s worth being precise about motive here, because “anti-scraping” is only one reason.
Bot mitigation. Vendors like Cloudflare, Akamai, and DataDome offer tiered responses as an alternative to outright blocking. A suspicious-but-not-confirmed request might get a JS challenge page, a reduced dataset, or a cached generic version instead of the personalized or full page. This buys the detection system more signal without tipping off the requester that they’ve been flagged, and it’s cheaper on their infrastructure than serving full dynamic content to every automated hit.
Legitimate performance and personalization, not detection. A lot of “different content” has nothing to do with bots at all. CDN edge caches serve a generic cached page to first-time or cookie-less visitors and a personalized one to sessions with state. A/B testing frameworks split traffic by cohort. Geo-routing serves a different page by inferred location, and datacenter proxy IPs often geo-resolve to the wrong place or to a hosting provider’s address block rather than a residential one. If your scraper’s IP resolves to “Ashburn, Virginia, hosting provider” instead of the city an actual user would show up in, you may get the generic template for reasons that have nothing to do with fingerprinting.
Rendering path differences. A single-page app that hydrates content via a client-side fetch call will show a scraper the loading skeleton and show a browser the populated page, purely because the scraper never ran the JavaScript that makes the second request. This is the most commonly misdiagnosed case: it looks exactly like deliberate discrimination but the site isn’t doing anything targeted at all.
Separating these three causes matters because the fix is completely different depending on which one you’re hitting.
How to tell which one you’re dealing with
Don’t guess. Diagnose it the same way you’d debug any other network issue: compare requests side by side.
- Capture the exact request your scraper sends. Log every header, in order, plus the TLS client if you’re setting one explicitly. Most scraping libraries expose this if you turn on debug logging.
- Capture the exact request a real browser sends to the same URL, using browser devtools’ network tab with “copy as cURL” or similar. Diff the two header sets.
- Check whether the response references client-side data fetches. View the raw HTML source (not the rendered DOM) for the scraper’s response. If the content you want lives inside a
<script>tag as JSON, or the DOM is mostly empty divs with data-attributes, that’s a rendering-path issue, not a detection issue. - Check the IP’s reverse DNS and ASN. If it resolves to a known hosting or proxy provider, geo-based and reputation-based serving may explain content differences that have nothing to do with headers or fingerprints.
- Look at response headers, not just body content.
Vary,Cache-Control,Set-Cookie, and anyX-custom headers from a CDN can tell you whether you got a cache hit versus a dynamically generated response. AVary: User-Agentheader is a direct admission that content is branching on that field.
This kind of side-by-side diff usually narrows the cause to one of the three buckets above within a few requests, rather than leaving you swapping settings at random.
What responsible operators actually do about it
None of this is a checklist for evading detection, and there isn’t a version of that checklist we’d publish here even if you asked, because “make requests indistinguishable from a browser” runs straight into a site’s terms of service and, depending on what’s being accessed, real legal exposure. What’s worth saying plainly instead:
Operators running production pipelines treat consistent, complete headers and realistic TLS behavior as baseline request hygiene, the same way they’d treat setting a correct Content-Type on a POST. If your target renders content client-side, a headless browser is the correct tool for the job, not a workaround, because you actually need the JavaScript to run to get the data that exists. Residential or mobile-carrier proxy infrastructure exists because IP reputation and geo-resolution are real signals that affect content and rate limits, not because it makes you undetectable, and no proxy vendor that’s honest with you will claim otherwise.
Beyond the technical layer, the boundary that actually matters is what the site’s terms of service and robots.txt say, and whether what you’re collecting is public, non-personal, and not paywalled. A differential response is sometimes the site telling you, in the only language HTTP has, that it doesn’t want this kind of traffic. Diagnosing why it’s happening is good engineering. Working around it just because you can isn’t the same decision as being allowed to.
If you’re building a scraping pipeline and want the header, TLS, and rendering-path diagnostics above wired into your monitoring instead of done by hand each time, that’s the kind of infrastructure problem we write about and build tooling around.
Check out more breakdowns like this one on the Data Research Tools home page.
Get new guides and videos first — join the Telegram channel.
Leave a Reply