What Scrapy’s Default Settings Actually Cost You

Nobody reads default_settings.py before their first crawl. You run scrapy startproject, write a spider, hit go, and it works. That’s the whole point of a framework: it gets you moving before you understand what you just turned on. The problem shows up later, usually in production, usually as a support ticket that says “the crawler stalled” or “we’re getting banned everywhere” with no obvious cause.

Every one of those tickets traces back to a setting that was never touched. Scrapy’s defaults aren’t wrong, they’re tuned for getting a first crawl running without configuration. Once you’re running a real pipeline against real sites, every one of those defaults has a cost attached, and most of them are invisible until something breaks.

The user agent that announces itself

The stock USER_AGENT value is Scrapy/VERSION (+https://scrapy.org). It’s built from the installed package version at import time, which means it’s not just identifying you as an automated client, it’s telling the target exactly which library and roughly which release you’re running, with a link to the project homepage attached.

This isn’t a stealth failure, it’s not meant to be stealthy at all. It’s a courtesy string so site operators know what’s hitting their logs and can find the framework’s docs if they want to set a crawl policy. But if you ship it to production unmodified, you’ve told every log-reading system on the other end what you are before it even has to look at your request pattern. The cost here isn’t dramatic, it’s just wasted effort: any WAF rule, rate limiter, or analyst doing a five-minute log review gets an easy signal for free. Setting a real, project-specific UA string is one of the first things any production config changes, and it should be paired with the rest of your request fingerprint, not treated as the whole fix.

Concurrency with no pacing

CONCURRENT_REQUESTS defaults to 16 globally, CONCURRENT_REQUESTS_PER_DOMAIN defaults to 8, and DOWNLOAD_DELAY defaults to 0. AUTOTHROTTLE_ENABLED is off. Put together, that’s a spider that will happily open eight simultaneous connections to a single domain with zero delay between requests, the moment it has eight URLs queued for that domain.

For a tutorial site or your own test server, that’s fine. Against a real production target, that request pattern is a burst, and bursts are exactly the shape that rate limiters and basic bot-detection heuristics are built to catch. It doesn’t matter how good your proxy pool is if the timing of the requests looks like a script instead of a browser session. A person clicking through pages doesn’t fire eight requests in the same tenth of a second.

The cost compounds because of what happens next: RETRY_TIMES defaults to 2, and Scrapy retries on a defined set of HTTP status codes including 429 and 503, which are the exact codes a rate limiter returns. So the default behavior when you get throttled is to immediately retry into the same throttle, twice, still with no added delay unless you’ve turned on AUTOTHROTTLE_ENABLED or set your own DOWNLOAD_DELAY. That’s compute and IP reputation spent making the problem worse, not better. In a farm running against multiple targets, an unthrottled spider hitting one flaky domain can burn through retry budget and proxy trust for the whole run.

The robots.txt setting that depends on how you started

This is the one that catches experienced people, not beginners. The actual library default for ROBOTSTXT_OBEY in scrapy/settings/default_settings.py is False. But the settings.py file that scrapy startproject generates for you sets it to True. So depending on whether your spider was born from the CLI template or built by importing Scrapy directly into an existing codebase, orchestration script, or Docker image, you can end up with opposite behavior for the exact same “default” setting.

This matters operationally, not just legally. Robots.txt is a policy signal, and plenty of teams have a compliance requirement to respect it regardless of what’s technically enforceable. If your production spiders were bootstrapped outside the standard template, either by copying an old settings file, building a custom base spider class, or running Scrapy as a library inside another framework, it’s worth explicitly checking which value you actually have rather than assuming the CLI default followed you. This is a five-minute check that avoids a genuinely awkward conversation with legal or with a partner site later.

Cookies you didn’t ask for

COOKIES_ENABLED defaults to True. Scrapy will accept and replay cookies per spider by default via its cookie middleware, which is usually what you want for anything involving login state or session-based pagination. But it also means your crawl is building session continuity you may not have intended, and that continuity is itself a fingerprinting surface. A session that persists across thousands of requests to the same domain, at machine-precision timing, with a consistent (or absent) cookie jar, is a much stronger and more stable identifier than IP alone. If you don’t need session state for a given spider, turning cookies off removes one more axis a detection system can use to stitch your requests together across proxy rotations.

The debug console left listening

TELNETCONSOLE_ENABLED is True by default, giving you a live telnet shell into the running process for inspecting the engine, stats, and queues, bound to 127.0.0.1 by default. That’s a genuinely useful debugging tool. The cost shows up in containerized deployments, where “localhost” doesn’t mean what it means on a bare-metal box. Depending on your network mode, a container’s loopback interface can end up reachable from other places you didn’t intend, and a debug console into a running scraping process is not something you want discoverable. It’s worth confirming explicitly whether it’s needed in your production images, rather than inheriting it because nobody revisited the setting after the container config changed hands.

DNS caching in a rotating-proxy world

DNSCACHE_ENABLED defaults to True with a cache size of 10,000 entries. That’s a sensible default for a single long crawl against a fixed target: don’t re-resolve names you’ve already resolved. But in an infrastructure where proxies rotate and requests may resolve DNS differently depending on which upstream or resolver a proxy is using, a long-lived, unbounded-feeling cache inside the Scrapy process can quietly serve stale resolution data across a run that spans hours. It rarely causes an outright failure, it shows up as odd intermittent behavior that’s hard to reproduce, because the caching layer is invisible unless you go looking for it.

What actually changes for production

None of this means Scrapy’s defaults are badly chosen. They’re chosen correctly for what they’re for: get a new user to a working crawl in ten minutes without reading documentation. The cost only appears when a tutorial config gets promoted to a production pipeline without anyone doing a settings audit in between.

The fix isn’t a single flag, it’s a checklist: a real, maintained User-Agent tied to your actual project, deliberate concurrency and delay values matched to the target rather than the framework ceiling, AUTOTHROTTLE_ENABLED turned on so retry storms don’t compound rate-limit hits, an explicit and documented decision on ROBOTSTXT_OBEY rather than an inherited one, cookies scoped to spiders that actually need session state, and the telnet console reviewed against your actual container network model. Every one of these is a two-line settings change. The cost of skipping them is measured in wasted proxy budget, incomplete crawls, and debugging sessions that start with “it worked in dev.”

If you’re running scraping infrastructure at any real scale, treating the settings file as part of the codebase to review, not a one-time bootstrap step, is the difference between a pipeline that degrades gracefully and one that quietly bleeds retries and reputation until someone notices the data stopped coming in.

Want more of this kind of infrastructure-level breakdown? Head back to the Data Research Tools home page for our pipeline orchestration, proxy infrastructure, and bot-detection explainers.

Get new guides and videos first — join the Telegram channel.

Comments

Leave a Reply

Your email address will not be published. Required fields are marked *