Your cart is currently empty!
Robots.txt and the legal line for scrapers
A text file that everyone treats as law and almost no one enforces
Every scraping team runs into robots.txt in the first week. Someone checks it, sees a Disallow: /api/ line, and asks whether that means scraping the API is illegal. The honest answer is that robots.txt is not law. It is a voluntary signaling protocol, and what it signals is intent, not permission or prohibition in a legal sense. Understanding what the file actually does, mechanically, is the only way to reason clearly about the legal questions that come after it.
What robots.txt actually is
The Robots Exclusion Protocol dates back to 1994. A site places a plain text file at /robots.txt with rules like:
User-agent: *
Disallow: /admin/
Crawl-delay: 2
Sitemap: https://example.com/sitemap.xml
Nothing on the server enforces this file. There’s no gate, no token check, no authentication layer tied to it. A crawler reads the file voluntarily, parses the User-agent block that matches its own identifier, and then chooses whether to honor the Disallow and Crawl-delay directives. Googlebot, Bingbot, and most reputable crawlers do honor it, because ignoring it can get a crawler’s IP ranges blocked and because search engines have a business reason to stay on good terms with site owners. A scraper you write yourself has no such obligation baked in. Whether it respects robots.txt depends entirely on whether you tell it to.
That’s the whole mechanism. It is a courtesy notice, not a firewall. Any operator running a real pipeline should already know this, because we build our own fetchers and none of them get robots.txt enforcement for free from the OS or the HTTP stack. We have to read the file and implement the check ourselves, the same way a browser doesn’t check it either.
Where the actual legal exposure comes from
Because robots.txt isn’t enforced technically, it also isn’t the thing courts have leaned on most heavily when scraping disputes go to litigation. Two bodies of law matter more:
The Computer Fraud and Abuse Act (CFAA) is the US federal statute most often cited in scraping cases. It criminalizes accessing a computer “without authorization” or “exceeding authorized access.” For years, defendants argued that violating a site’s terms of service, or scraping past a robots.txt disallow, counted as “without authorization.” Courts have narrowed that reading substantially. In Van Buren v. United States (2021), the Supreme Court held that “exceeds authorized access” applies to accessing off-limits areas of a system you’re otherwise permitted to use, not to misusing access you already have for a purpose the owner dislikes. In hiQ Labs v. LinkedIn (9th Circuit, most recently affirmed in 2022 after remand), the court held that scraping data that is publicly accessible without logging in generally does not violate the CFAA, because there’s no access barrier to circumvent in the first place. Robots.txt wasn’t the deciding factor in either case; the presence or absence of an actual technical access barrier was.
Contract and terms-of-service law is the other track, and it runs independently of the CFAA. A site’s terms of service can prohibit automated access even for public pages. If you create an account, or even just use the site, under terms that ban scraping, a court can treat that as a breach of contract claim rather than a computer-crime claim. This is a lower bar for a plaintiff to clear than CFAA liability, and it doesn’t require proving anything about robots.txt at all. Some site owners have also succeeded with trespass-to-chattels theories, arguing that scraping traffic imposed a measurable burden on their servers.
So the practical picture is: robots.txt is evidence of the site owner’s stated wishes, which can be relevant context in a dispute, but the actual liability usually turns on whether there was an access barrier to bypass (CFAA), what the terms of service said (contract law), and what kind of data was collected (privacy statutes like GDPR or CCPA apply regardless of what any robots.txt file says, because they’re about the nature of the data, not the method of collection).
What this means for how we build
None of this is a green light to disregard robots.txt because it “isn’t technically law.” In production infrastructure, respecting it is good engineering practice independent of the legal question:
- Crawl-delay and rate limits protect your own infrastructure too. A crawler that ignores a site’s stated crawl-delay is more likely to get flagged by the target’s bot-detection layer, burn through proxy pool reputation faster, and trigger IP-range blocks that cost you working exit nodes. Respecting the stated delay is cheaper than fighting the fallout.
- Disallow directives often mark the exact paths a site’s detection systems watch most closely. A
/Disallow: /search-results/line is frequently a hint about where rate-based or fingerprint-based defenses are tuned tightest. Site operators use crawl-delay and disallow rules as one signal among several (alongside request headers, TLS fingerprints, and behavioral patterns) to separate expected crawler traffic from unexpected load. - The file tells you what the site owner considers acceptable, which matters for the contract-law exposure even when it doesn’t resolve the CFAA question. If a dispute ever gets to a demand letter or a court filing, a scraper that ignored an explicit, published Disallow directive on a site whose terms also prohibited scraping is in a materially worse position than one that stayed within stated bounds.
The parts robots.txt says nothing about
Robots.txt has no concept of personal data, and it is not a privacy control. A page can be fully allowed by robots.txt and still contain personal data whose collection, storage, or resale is regulated separately under GDPR, CCPA, or similar statutes depending on jurisdiction and the nature of the data subjects. Scraping something robots.txt permits doesn’t clear that separate legal question, and scraping something it disallows doesn’t automatically create liability under privacy law either. These are different axes entirely, and treating “robots.txt allowed it” as a compliance checkbox for personal data is a mistake we see teams make regularly.
Robots.txt also says nothing about paywalled or authenticated content. Whatever a site allows crawlers to see through robots.txt applies to publicly reachable pages. Content sitting behind a login or a paywall is a different legal and technical situation, generally involving actual access controls, which is exactly the kind of barrier the CFAA and hiQ line of cases treat as significant.
The bottom line for anyone running this infrastructure
Robots.txt is a directive, not a lock. It tells you what a site owner has published as their preference, and courts have generally treated it as one data point among several rather than a dispositive legal boundary. The stronger, more consistent legal lines run through whether there was an actual access barrier, what a site’s terms of service say, and what kind of data is being collected. None of this is a substitute for legal advice about a specific use case and specific jurisdiction; it’s the operating context worth understanding before you decide how a crawler behaves.
If you’re building scraping infrastructure and want the engineering side, from proxy pool design to how bot detection actually evaluates a request, that’s what we cover here at Data Research Tools.
Get new guides and videos first — join the Telegram channel.
Leave a Reply