Your cart is currently empty!
Error taxonomy: what each status code actually means for a scraper
Most scraping pipelines treat every non-200 response the same way: log it, retry it, move on. That works until volume goes up and the retry logic starts making things worse. A 429 and a 500 are not the same problem, and treating them identically burns proxy budget on one and gives up too early on the other. Running proxy infrastructure at scale means building an actual taxonomy, not just a catch block.
This is the breakdown we use operationally. It’s not a list of codes with dictionary definitions. It’s what each one tells you about where the failure is happening, and what response actually makes sense.
Why the taxonomy matters more than the retry count
A generic scraper retries on failure and gives up after N attempts. That’s fine for a hobby script pulling a few hundred pages a day. At production volume, undifferentiated retries do two bad things: they hammer a target that’s already rate limiting you, which pushes you further into a block, and they waste proxy sessions retrying something that will never succeed no matter how many times you ask, like a 404 on a page that was actually removed.
The fix is classifying every response into one of a small number of buckets before deciding what to do with it: succeeded, needs a different identity, needs to slow down, target-side failure, or gone. Each bucket has a different correct action. Get the classification wrong and your pipeline either burns through proxies for nothing or gives up on pages that would have worked on the next try.
2xx doesn’t always mean success
A 200 status code means the server sent a response body. It does not mean the response body is the page you wanted. Sites serving a soft block, a CAPTCHA challenge page, or an interstitial “verify you’re human” screen frequently return 200. The HTTP layer is happy. The content layer is empty or wrong.
This is the single most common blind spot in scraper monitoring dashboards that only track status codes. If your success metric is “percentage of 200s,” you can be blocked on 40% of requests and your dashboard will show green. The only real fix is validating the payload, not just the status line: check for expected selectors, a minimum content length, or a known marker string before counting a request as successful. Status code and content validity are two separate checks and both need to pass.
3xx: the site telling you it moved, or telling you something else
Redirects are usually benign. A 301 on a product page that got a new URL slug, a 302 during a login flow, these are normal site behavior and following them is correct.
The pattern worth watching is a redirect chain that always lands on the same destination regardless of the requested URL, often a login wall, a “please enable JavaScript” page, or a generic error page. That’s a redirect being used as a soft block. The distinction matters for your taxonomy: a 301 to a plausible new URL is a success path, a redirect that funnels everything to one fixed landing page is a block and should be logged as one, not silently followed and counted as a completed fetch.
401 and 403: two different problems that look similar
401 means the request needed credentials and didn’t have valid ones. In scraping contexts this shows up on APIs and authenticated endpoints, and it’s usually a configuration problem on your end: expired token, wrong header, session cookie that didn’t get carried over.
403 is broader and more common on scraping targets. It can mean an IP-based block, a fingerprint-based block, a WAF rule matching the request pattern, or geographic restriction. Sites make these decisions using signals like request rate from a single IP, TLS handshake fingerprint, header ordering and completeness, and behavioral patterns across a session. This is defensive infrastructure working as intended from the site’s side, and it’s worth understanding it that way rather than as a random obstacle: the 403 is a decision, not a glitch, and it usually means something about the request pattern, not just the destination, tripped a rule.
Operationally, a 403 should route differently than a 401. A 401 means fix the request. A 403 means the identity making the request, the IP and its associated fingerprint, has been flagged, and repeating the identical request from the same identity is very unlikely to change the outcome.
404: gone, moved, or never existed, and each needs a different response
A 404 on a URL you scraped successfully last week usually means the page was actually removed or the site restructured its URL scheme. A 404 on a URL you constructed yourself, say by incrementing an ID, might mean you’ve walked past the end of a valid range, not that anything is wrong.
The mistake we see most often is treating 404 as a transient error and retrying it. It isn’t transient. Retrying a 404 ten times wastes ten requests confirming what the first one already told you. The correct action is to log it once, mark that URL as resolved-not-found, and move on. If 404 rate on previously-valid URLs suddenly spikes across a whole domain, that’s a signal the site restructured, and it’s worth flagging for a human to check the URL pattern rather than continuing to auto-retry.
429: the most honest status code in the list
429 (Too Many Requests) is the site explicitly telling you the rate is the problem, not the identity, not the content of the request. It’s the clearest signal in the whole taxonomy because it names the mechanism.
The correct response to a 429 is to slow down, and specifically to respect a Retry-After header if the site sends one. Rotating to a different proxy immediately after a 429 without changing pace just moves the same rate-limiting pressure to a new IP, which is treating a pacing problem as an identity problem. Some pipelines do both, easing the per-identity rate and distributing load across a pool, but the rate change is the part that actually addresses what the 429 said. A pipeline that only ever rotates IPs on 429 and never adjusts request pacing is optimizing the wrong variable.
5xx: figuring out if it’s them or you
500, 502, 503, and 504 all indicate the failure is server-side, but they’re not equally informative. A 500 means the origin server hit an unhandled error, often unrelated to you specifically. A 502 or 504 usually means something in front of the origin, a load balancer or reverse proxy, couldn’t get a valid response from the backend in time. A 503 is frequently deliberate: “service unavailable” is also how some infrastructure signals load shedding or maintenance, and on sites with bot defenses it can appear as a temporary block state rather than a genuine outage.
The practical test is whether the 5xx rate correlates with your own request volume or shows up independent of it. If a small, well-paced scraper gets 503s at the same rate as a request storm would, the site’s front end is treating your traffic as excessive load, whether or not that’s the actual cause. If 5xx rates track a known site outage or deploy window, it’s genuinely their problem and a longer backoff with a handful of retries is the right move.
The failures with no status code at all
A meaningful share of scraping failures never get a status code because the connection didn’t complete: DNS resolution failures, TCP connection refused, TLS handshake failures, and read timeouts where the connection opened but no response came back in time.
These deserve their own bucket because they point at infrastructure, not content. A connection refused from every request through one proxy but not others points at that proxy or that egress IP, not the target site. A TLS handshake failure that only happens on certain proxy exit nodes usually means an outdated or misconfigured client stack on that node. Lumping these in with 5xx server errors hides infrastructure problems on your own side of the connection.
Putting it together
A working taxonomy sorts every outcome into a small number of buckets: validated success (200 plus content check passed), soft block (200 with failed content check, or a redirect loop to a fixed destination), rate limited (429, respond by slowing down), identity blocked (403, respond by rotating identity, not just retrying), resolved-not-found (404, log once, don’t retry), target-side failure (5xx, backoff and limited retry), and connection-level failure (no status code, investigate the proxy or network path). Each bucket maps to one clear action. The value isn’t in the categories themselves, it’s in refusing to let a single generic retry loop handle all seven of these differently-shaped problems the same way.
If you’re building or debugging a scraping pipeline and want more on the infrastructure side of this, proxy behavior, fingerprinting, and how detection systems actually work, we write about it regularly at Data Research Tools.
Get new guides and videos first — join the Telegram channel.
Leave a Reply