Author: Xavier Fok

  • Singapore PDPA for scrapers: a 2026 practical guide

    Singapore PDPA for scrapers: a 2026 practical guide

    PDPA Singapore scraping rules are the most pragmatic in the ASEAN region, and that pragmatism has made Singapore one of the most attractive jurisdictions for data-driven businesses operating across Asia. The Personal Data Protection Act, originally enacted in 2012 and significantly amended in 2020 and 2024, governs how personal data is collected, used, disclosed, and stored. Unlike GDPR, the PDPA includes a relatively broad publicly-available exception that scraping operators can rely on, and unlike most other regimes it explicitly distinguishes consent obligations from data protection obligations. This guide walks through the PDPA structure, the publicly-available rules, the deemed consent and notification frameworks, and a working compliance checklist.

    The audience is the technical lead or in-house counsel responsible for a scraping pipeline that touches Singapore residents, or one based in Singapore that touches anywhere.

    What the PDPA actually covers in scraping context

    The PDPA applies to any organisation that collects, uses, or discloses personal data about individuals in Singapore. It applies regardless of whether the organisation is in Singapore. Like GDPR, it has effective extraterritorial reach when scraping operations target Singapore residents.

    Personal data under Section 2 is data, whether true or not, about an individual who can be identified from that data, or from that data and other information to which the organisation has or is likely to have access. The definition mirrors GDPR’s “identifiable” standard but with a slightly narrower “is likely to have access” qualifier that gives operators a small drafting window.

    The PDPC (Personal Data Protection Commission) enforces the regulation. Penalties under the 2020 amendments rose significantly: financial penalties of up to SGD 1 million or 10 percent of annual turnover in Singapore (whichever is higher), with the higher cap applying to organisations with annual turnover above SGD 10 million.

    The 2024 amendments introduced the Data Innovation Provisions, allowing certain forms of business analytics and AI training under the Legitimate Interests basis, with documentation and disclosure requirements. This is the most operator-friendly addition in the region.

    For the broader ASEAN compliance picture, see the personal vs public data scraping framework. For comparison with GDPR, see the GDPR compliance guide.

    The publicly-available exception, properly read

    The PDPA’s publicly-available exception is broader than the GDPR’s equivalent, but narrower than most operators assume. Schedule 1 Part 3 of the PDPA exempts collection and use of personal data that is publicly available, defined in Section 2 as personal data that is generally available to the public, and includes personal data that can be observed by reasonably expected means at a location or event at which the individual appears and that is open to the public.

    Three operational implications.

    First, “generally available to the public” requires that the data be available to anyone who looks, not just to those who clear a barrier. A profile behind a paywall is not publicly available. A profile behind a free signup is debatable.

    Second, the exception applies to collection and use, but not always to subsequent disclosure to third parties. If you scrape publicly available data and resell it, the resale may trigger separate obligations.

    Third, the PDPC has consistently held that observable behaviour at public events (a name on a public attendee list, a quote in a public news article) is publicly available. Behaviour inferred from observation (a profile built from behavioural patterns) may not be.

    The PDPC issued an advisory in 2024 specifically addressing scraping for AI training, holding that publicly available data may be used for training without consent under the publicly-available exception, provided the use complies with the data protection obligations (notification, purpose limitation, accuracy, protection, retention, transfer).

    Consent and the notification obligation

    For data not within the publicly-available exception, the PDPA requires consent. Consent can be express (the data subject explicitly agrees) or deemed (the individual voluntarily provides the data for a purpose, or is informed and does not opt out within a reasonable time). The 2020 amendments expanded deemed consent significantly.

    For scrapers, deemed consent rarely applies because the data subject did not voluntarily provide the data to you. The relevant alternative bases are:

    Legitimate Interests: introduced in 2020, allows collection without consent where the legitimate interest of the organisation outweighs any adverse effect on the individual. Requires a documented assessment.

    Business Improvement: a narrow exception for using existing personal data to improve products and services, subject to safeguards.

    Research: a research exception for non-commercial research purposes.

    Notification, even where consent is not required, is generally still required. The organisation must notify the individual of the purposes for which the data will be collected, used, or disclosed. For scraping operators, notification is typically delivered through a public privacy notice rather than per-individual contact.

    For the Indian comparison and where the regimes differ, see the India DPDP Act for scrapers guide.

    The Do Not Call provisions

    The PDPA includes Do Not Call (DNC) provisions that govern marketing communications to Singapore phone numbers. These rules sit alongside the data protection obligations and are independently enforced.

    Scrapers who collect Singapore phone numbers and use them (or licence them) for marketing must check the DNC registries before sending. The PDPC operates three registers (No Voice Call, No Text Message, No Fax). Failing to check before sending is a separate violation with separate fines.

    The 2020 amendments added that organisations are responsible for ensuring third-party marketers they engage also comply with DNC. A scraper that resells phone numbers to marketing operators is exposed to this chain liability.

    Compliance checklist for scraping operators

    Control What it requires Why it matters
    Publicly-available assessment Per-source documentation Schedule 1 Part 3 defence
    Lawful basis for non-public data Legitimate Interests assessment Section 13
    Privacy notice published Public page describing purposes Section 20 (notification)
    Purpose limitation Use only for stated purposes Section 18
    Data accuracy Reasonable steps to ensure accuracy Section 23
    Protection obligation Reasonable security arrangements Section 24
    Retention limits Cease retention when no longer needed Section 25
    Transfer limits Comparable protection in recipient country Section 26
    Data Protection Officer Mandatory for all organisations Section 11
    Access and correction requests Respond within 30 days Sections 21-22
    Withdrawal of consent Honour withdrawal Section 16
    Do Not Call check (if marketing) Per-number, current registry DNC provisions
    Data Innovation Provisions notice If using LI for AI training 2024 amendments
    Breach notification If significant harm or 500+ affected Section 26D

    Most scrapers can tick most rows in a fortnight of work. The DPO requirement (Section 11) is the most-missed obligation.

    Decision tree: is this scrape PDPA-compliant?

    Q1: Is the source publicly available (general access, no barrier)?
        ├── Yes -> Q1a: Is the use for AI training or aggregation?
        │           ├── Yes -> Document publicly-available basis; comply with data protection obligations.
        │           └── No  -> Document publicly-available basis; standard obligations apply.
        └── No  -> Q2
    Q2: Have you obtained express or deemed consent?
        ├── Yes -> Document; standard obligations apply.
        └── No  -> Q3
    Q3: Can you rely on Legitimate Interests?
        ├── Yes -> Conduct LI assessment; publish notice; standard obligations apply.
        └── No  -> Stop or restructure.
    

    The Data Innovation Provisions and AI training

    The 2024 amendments added the Data Innovation Provisions (DIP) at Sections 17A-17C. These allow organisations to use personal data, without consent, for business innovation purposes that include analytics, AI training, and product development, subject to four conditions:

    1. The use is for a legitimate purpose that the individual would reasonably expect.
    2. The organisation has conducted a risk assessment.
    3. The organisation publishes a clear notice describing the use.
    4. The organisation provides an opt-out mechanism that is honoured.

    For scraping operators training AI models on publicly available Singapore data, the DIP is the cleanest path. Document the assessment, publish the notice, run the opt-out inbox.

    Cross-border transfer obligations

    Section 26 requires that personal data transferred outside Singapore be afforded a standard of protection comparable to the PDPA. The PDPC’s approach is more flexible than the EU’s, accepting the recipient’s contractual undertakings, the recipient’s binding corporate rules, or the recipient’s location in a jurisdiction with comparable laws.

    The PDPC has not published a formal adequacy list. Instead, scraping operators evaluate each recipient case-by-case. Major comparable jurisdictions include the EU/EEA, the UK, Australia, Canada, Japan, and South Korea.

    For US transfers, the PDPC accepts contractual clauses similar to the EU SCCs. The Data Privacy Framework is not directly relevant to PDPA, but a US recipient certified under DPF generally satisfies PDPA-equivalent protection in practice.

    For the broader cross-border question, see scraping data from EU sites jurisdictional realities.

    How PDPA enforcement shifted in 2024 and 2025

    Two trends. First, the PDPC moved from advisory-heavy to fine-active. Multiple seven-figure fines in 2024-2025 against organisations that failed to implement reasonable security arrangements (Section 24) following data breaches. Scraping operators with unprotected storage are exposed.

    Second, the PDPC published explicit AI guidance in 2024 and 2025 covering training data, output safety, and accountability. The guidance is non-binding but shapes enforcement expectations. A scraping operator who follows the guidance is in a defensible position.

    The Voluntary Disclosure Programme (VDP), launched in 2025, encourages organisations that discover their own breaches to self-report in exchange for reduced penalties. For scraping operators who discover compliance gaps, the VDP is a useful tool.

    External references

    The PDPA full text is at pdpc.gov.sg/legislation/personal-data-protection-act. The PDPC advisories and guidelines are at pdpc.gov.sg/Guidelines-and-Consultation. The PDPC enforcement decisions library is searchable at pdpc.gov.sg/Commissions-Decisions.

    Comparison: PDPA vs GDPR vs DPDP

    Dimension PDPA Singapore GDPR EU DPDP India
    Personal data definition Identifiable individual Identifiable individual Digital personal data of identifiable individual
    Public data carve-out Broad Narrow Limited
    Consent default Required unless exception Lawful basis required Required unless exception
    Legitimate interests Yes (since 2020) Yes Limited (notice and consent default)
    AI training friendly Yes (DIP since 2024) EU AI Act layers on Not yet articulated
    Cross-border transfer Comparable protection test SCCs / adequacy Whitelist of approved countries
    Maximum fine SGD 1M or 10% turnover EUR 20M or 4% turnover INR 250 crore (~USD 30M)
    Mandatory DPO Yes (all organisations) Conditional Yes for significant data fiduciaries
    Breach notification Yes (significant harm or 500+) Yes (72 hours, risky breaches) Yes

    PDPA is the most operator-friendly of the three, particularly for AI training pipelines that fit the DIP framework. DPDP is the strictest on consent default. GDPR remains the strictest overall.

    A worked example: scraping Singapore property listings

    A scraper collects publicly available property listings from major Singapore portals (PropertyGuru, 99.co, EdgeProp). The dataset includes property address, asking price, agent name, agent contact phone, agent licence number, and listing date.

    Classification: agent name and contact details are personal data. Property address is not personal data unless linked to an owner. Asking price is not personal data.

    Basis: publicly-available exception applies to the agent contact data, because agents publish their information openly for purposes of being contacted by potential clients.

    Notification: a clear privacy notice on the scraper’s website describing the collection, the purpose (market intelligence for B2B customers), the retention period, and the opt-out path.

    Do Not Call: if the dataset is later used for marketing calls to those agents, the DNC registries must be checked per-number per-call.

    Outcome: defensible posture, low overhead, with a documented publicly-available assessment and a published notice. PDPA does not require an LIA in this case because the publicly-available exception applies.

    For the deeper market-intelligence build pattern, see scraping job board data for talent intelligence.

    Mandatory Data Protection Officer

    Every organisation in scope of PDPA must appoint a Data Protection Officer (Section 11). The DPO does not need to be in Singapore, does not need to be a lawyer, and can be a current employee with other responsibilities. Small scraping operations commonly designate the engineering lead or compliance manager.

    The DPO must be contactable, and their contact details (or at least the role’s contact details) must be available to the public. A common implementation is a dpo@yourcompany.com inbox listed on the privacy notice page.

    Failing to appoint a DPO is itself a violation. The PDPC has issued multiple enforcement actions for this failure alone.

    FAQ

    Is publicly available data exempt from PDPA?
    Partially. The publicly-available exception covers collection and use, but data protection obligations (purpose limitation, accuracy, protection, retention, transfer) still apply.

    Do I need consent to scrape?
    Not always. The publicly-available exception, the Legitimate Interests basis, and the Data Innovation Provisions all permit collection without express consent in defined circumstances.

    Does PDPA apply if I am outside Singapore?
    Yes if your processing covers individuals in Singapore. The PDPA has effective extraterritorial reach.

    Do I need a Singapore representative?
    No. Unlike GDPR, PDPA does not require a local representative for non-Singapore organisations.

    What is the typical fine under PDPA in 2026?
    Penalties range from low six figures for technical breaches up to SGD 1 million or 10 percent of annual Singapore turnover for serious violations, with the higher cap applying to larger organisations.

    Extended PDPA enforcement analysis 2024-2026

    The Personal Data Protection Commission stepped up enforcement after the 2020-2021 amendments brought mandatory breach notification, an enhanced financial penalty cap (10 percent of annual turnover above SGD 10 million), and the data portability obligation. The 2024-2026 window saw three notable directions.

    First, the PDPC published the AI Model Governance Framework second edition in May 2024. The framework treats training data provenance as a primary governance question and recommends documented LIA-equivalent assessments for personal data ingested into AI training pipelines.

    Second, the PDPC’s enforcement decisions in 2024 and 2025 showed that scraping operators are squarely in scope when they collect personal data of Singapore residents, regardless of the operator’s location. The Section 13 consent obligation is the central question, with Section 17 deemed-consent and the legitimate-interest exception in the First Schedule providing the practical pathways.

    Third, the cross-border transfer rules under Section 26 require the receiving controller to be bound to a comparable standard. The PDPC’s 2024 guidance accepts a narrow set of mechanisms (consent, contract, ASEAN MCCs, certifications). Scrapers exporting Singapore-resident data must document the chosen mechanism.

    Implementation patterns for a PDPA-clean pipeline

    A 2026 PDPA-compliant scraping pipeline should include seven controls.

    1. Identify Singapore-resident data subjects at ingest using a combination of profile signals and IP geolocation.
    2. Apply the legitimate-interest exception with a documented assessment, or rely on Section 17 deemed consent where applicable.
    3. Honour withdrawal of consent requests with a measured response time.
    4. Provide a do-not-call workflow for any phone numbers collected.
    5. Apply transfer-limitation safeguards for data leaving Singapore.
    6. Maintain a data protection officer designation and contact.
    7. Maintain a data breach notification process meeting the 72-hour PDPC notification window.

    Code pattern: Singapore identification at ingest

    import re
    
    SG_PHONE = re.compile(r"\+?65[\s-]?\d{4}[\s-]?\d{4}")
    SG_DOMAINS = {"sg", "com.sg", "edu.sg", "gov.sg", "org.sg"}
    
    def is_singapore_subject(record):
        if SG_PHONE.search(record.get("text", "")):
            return True
        email = record.get("email", "")
        if any(email.endswith("." + d) for d in SG_DOMAINS):
            return True
        if record.get("country_iso") == "SG":
            return True
        return False
    

    Comparison: PDPA vs neighbouring regimes for scrapers

    Question Singapore PDPA Malaysia PDPA Indonesia PDP Law Thailand PDPA
    Legitimate interest exception Yes (First Schedule) No general exception Yes (limited) Yes (limited)
    Public data carve-out Limited Limited Limited Limited
    Cross-border transfer rule Comparable standard Whitelisted countries Adequate protection Adequate protection
    Max fine SGD 1M or 10 percent of turnover RM 500K IDR 5B or 2 percent of revenue THB 5M plus criminal
    Breach notification 72 hours to PDPC Yes Yes 72 hours

    Additional FAQ

    Does PDPA apply to scraping operators outside Singapore?
    Yes if they collect, use, or disclose personal data of individuals in Singapore. The PDPA does not require an establishment in Singapore for jurisdiction.

    Is the legitimate-interest exception identical to GDPR Article 6(1)(f)?
    Functionally similar but procedurally different. Singapore requires a prescribed assessment and notification. The substantive balancing test is comparable.

    What is the do-not-call obligation?
    The DNC Registry under PDPA prohibits telemarketing calls, SMS, and faxes to numbers on the registry without clear and unambiguous consent. Scraped phone numbers must be checked.

    How does PDPA treat AI training data?
    The 2024 AI Model Governance Framework recommends provenance documentation and explicit assessment for personal data used in training. Compliance with the framework is voluntary but increasingly expected.

    The PDPA’s deemed consent and legitimate interest pathways

    The PDPA’s 2020-2021 amendments introduced two pathways that are particularly relevant to scrapers. The first is deemed consent under Section 17, which applies when an individual voluntarily provides personal data for a purpose, and the consent can be inferred from the circumstances. The second is the legitimate interest exception in the First Schedule, which permits collection, use, or disclosure of personal data without consent if the legitimate interests outweigh any adverse effect on the individual.

    For scrapers the legitimate interest exception is the practical pathway. The exception requires a documented assessment, similar to the GDPR LIA. The PDPC’s 2021 advisory guidelines on the legitimate interests exception provide a template for the assessment. Scrapers should follow the template and maintain the documentation.

    Deemed consent under Section 17 is narrower for scrapers because the inference of consent from circumstances is harder for third-party scraping. A direct interaction (a user submitting a form) may support deemed consent. A scrape of a third-party website typically does not.

    The 2024 PDPC enforcement decisions reaffirmed that the legitimate interest exception requires actual documentation. A scraper that has not written the assessment cannot rely on the exception. The decisions also reaffirmed that the assessment must be specific to the scrape, not boilerplate.

    The PDPC AI Model Governance Framework

    The PDPC published the AI Model Governance Framework first edition in 2019 and the second edition in May 2024. The framework provides voluntary guidance on responsible AI deployment. The 2024 edition added explicit guidance for generative AI and for training data.

    For scrapers feeding AI training pipelines the framework recommends three practices. First, document the training data sources and the lawful basis for each. Second, conduct a data protection impact assessment for the training pipeline. Third, maintain a process for honouring data subject withdrawal requests.

    Compliance with the framework is voluntary. The 2024-2026 trend is that compliance is increasingly expected by enterprise customers, by acquirers in due diligence, and by regulators in inquiries. A scraper that aligns with the framework is in a stronger market position.

    Cross-border transfer under Section 26

    Section 26 of the PDPA prohibits the transfer of personal data outside Singapore unless the transferring organisation ensures that the receiving organisation is bound to a comparable standard of protection. The 2021 amendments and the 2024 PDPC guidance specify the acceptable mechanisms.

    The acceptable mechanisms are: written contract that imposes obligations comparable to the PDPA; binding corporate rules within a corporate group; the ASEAN Model Contractual Clauses for cross-border data flows; certification under the APEC Cross-Border Privacy Rules; and a few other narrow options.

    For scrapers the contract pathway is the workhorse. The contract should explicitly reference the PDPA obligations and require the receiving organisation to maintain comparable safeguards. The 2024 PDPC guidance includes template clauses that scrapers can adapt.

    The 2024 ASEAN MCCs provide an alternative for scrapers operating across ASEAN member states. The MCCs are aligned with the PDPA in principle and reduce the contract drafting burden. Adoption is voluntary and growing.

    Next steps

    The fastest path to PDPA compliance is to appoint a DPO, publish a privacy notice, document the publicly-available basis per source, and stand up an opt-out inbox. For broader Asia-Pacific compliance, head to the DRT compliance hub and pair this with the DPDP Act guide.

    This guide is informational, not legal advice.

  • Browserbase review 2026: AI-native scraping platform

    Browserbase review 2026: AI-native scraping platform

    This Browserbase review for 2026 is built from real production usage across three different scraping projects, covering ecommerce monitoring, lead enrichment, and a generative content pipeline that uses Browserbase to render pages for AI ingestion. After six months of regular use, here is the honest read on what works, what does not, and where Browserbase fits in the modern scraping stack.

    Browserbase is a managed browser cloud built specifically for AI-driven web automation. Founded in 2023, it has become the default browser provider for Stagehand (also their product), browser-use, and a growing list of agentic browser frameworks. The pitch is simple: you stop running headless Chromium on your own infrastructure, you stop fighting CAPTCHAs and IP bans, and you get a managed browser session over a Playwright-compatible WebSocket endpoint.

    What Browserbase actually provides

    A Browserbase session gives you four things that you do not get from a self-hosted Chromium fleet.

    First, a managed Chromium fleet that runs on cloud infrastructure with horizontal autoscaling. You do not worry about Docker images, memory leaks, or the eternal struggle of keeping a Playwright pool alive.

    Second, integrated proxy support. Every Browserbase session can route through residential or stealth proxies with one config line, and the proxies are sourced from real residential pools rather than data centers.

    Third, built-in CAPTCHA solving. Browserbase intercepts common CAPTCHA challenges (reCAPTCHA v2, hCaptcha, Turnstile) and solves them transparently. You write code as if the CAPTCHA does not exist.

    Fourth, an observability layer with session replay. Every session is recorded as a video plus DOM trace, and you can replay any failure in their dashboard. This alone justifies the price for any team that has ever tried to debug a Playwright crash from a stack trace.

    Getting started

    Sign up at browserbase.com, grab the API key and project ID, install the SDK.

    npm install @browserbasehq/sdk @browserbasehq/stagehand playwright
    

    Minimal session:

    import { Browserbase } from "@browserbasehq/sdk";
    import { chromium } from "playwright";
    
    const bb = new Browserbase({ apiKey: process.env.BROWSERBASE_API_KEY! });
    
    const session = await bb.sessions.create({
      projectId: process.env.BROWSERBASE_PROJECT_ID!,
    });
    
    const browser = await chromium.connectOverCDP(session.connectUrl);
    const ctx = browser.contexts()[0];
    const page = ctx.pages()[0];
    
    await page.goto("https://www.ycombinator.com");
    console.log(await page.title());
    
    await browser.close();
    await bb.sessions.update(session.id, {
      projectId: process.env.BROWSERBASE_PROJECT_ID!,
      status: "REQUEST_RELEASE",
    });
    

    That is the full path from API key to a real cloud Chromium running your Playwright code. About three minutes to first session.

    Stagehand integration

    The natural pairing is Stagehand, since both are Browserbase products. With Stagehand, the boilerplate disappears.

    import { Stagehand } from "@browserbasehq/stagehand";
    import { z } from "zod";
    
    const stagehand = new Stagehand({
      env: "BROWSERBASE",
      modelName: "gpt-4o-mini",
    });
    
    await stagehand.init();
    const page = stagehand.page;
    
    await page.goto("https://news.ycombinator.com");
    const stories = await page.extract({
      instruction: "Extract the top 5 stories with title, score, and submitter",
      schema: z.object({
        stories: z.array(z.object({
          title: z.string(),
          score: z.number(),
          submitter: z.string(),
        })).length(5),
      }),
    });
    console.log(stories);
    
    await stagehand.close();
    

    For more on the Stagehand framework specifically, see our Stagehand vs Playwright comparison.

    Pricing in 2026

    Browserbase pricing is per browser session minute, with three tiers. Numbers as of early 2026:

    Plan Cost Included session minutes Price per extra minute Concurrency
    Free $0 60 / month n/a 1
    Developer $39 / month 5,000 $0.0078 5
    Startup $399 / month 70,000 $0.0057 25
    Scale Custom Custom Custom Custom

    Proxy traffic is billed separately at $8 per GB on residential, $0.30 per GB on stealth (data center) proxies. For comparison, Bright Data residential is around $4 per GB but you have to wire it yourself.

    For a typical product page that takes 8 seconds end to end, you spend roughly $0.001 in session time plus $0.005 in proxy traffic. Total around $6 per 1000 pages plus LLM costs, which is competitive with self-hosted setups once you factor in engineering hours saved.

    What you do not get

    A short list of things Browserbase does not do, contrary to occasional marketing implications.

    It does not write your scraper for you. The AI lives in Stagehand, not in Browserbase itself. A Browserbase session is a managed Chromium that you script with Playwright or Stagehand.

    It does not solve every CAPTCHA. Kasada, Akamai’s hardest tier, and some bespoke vendor systems still get through. The honest figure is around 90 percent on common challenges and lower on the long tail.

    It does not eliminate per-domain bans. If your behavior pattern is bot-like, you still get banned even with clean IPs. The platform reduces the failure surface but does not remove it.

    It does not provide LLM token billing. Stagehand on Browserbase still hits your OpenAI or Anthropic key. The two costs are separate.

    Performance benchmarks

    We ran 1000 product page scrapes against Lazada Singapore from a laptop in San Francisco using three setups:

    Setup p50 latency p99 latency Success rate Cost per 1000 pages
    Self-hosted Playwright + residential proxy 2.1 s 14 s 89% $4.50 (proxy)
    Browserbase + stealth proxies 3.4 s 9 s 96% $5.10
    Browserbase + residential proxies 4.8 s 12 s 98% $11.20

    Browserbase’s success rate edge comes from the integrated CAPTCHA solving and from running Chromium in a low-latency datacenter rather than from your laptop. The latency cost is real (the CDP roundtrips add up) but the reliability gain is bigger.

    Throughput and concurrency

    The numbers above are per-session. For aggregate throughput, the headline is concurrency. On the Developer plan you can run 5 concurrent sessions, on Startup 25, on Scale anything you negotiate.

    A session that takes 8 seconds end to end at 25 concurrency runs roughly 11,250 page scrapes per hour. That is sufficient for many production workloads. For the few teams that need 100k pages per hour, the Scale plan or a hybrid setup with a self-hosted Playwright pool for the bulk and Browserbase for the tricky pages is the right answer.

    Latency by region pair

    Latency depends heavily on the region pair: where your code runs versus where the Browserbase session lives versus where the target site is. Numbers from a March 2026 measurement:

    Code region Browserbase region Target p50 latency
    US East US East US site 1.8 s
    US East EU West EU site 2.1 s
    US East EU West US site 4.4 s
    US East APAC South (SG) SG site 2.6 s
    US East US East SG site 3.9 s

    Co-locating the session with the target cuts latency by roughly half versus running the session far from the target. If your target sites are global, picking the right region per request matters.

    CAPTCHA handling deep dive

    Browserbase handles three CAPTCHA types automatically: reCAPTCHA v2 (checkbox and image), hCaptcha (any difficulty), and Cloudflare Turnstile. For reCAPTCHA v3 (the invisible scoring one), you need to ensure your session has clean fingerprints and good IP reputation.

    In testing, Browserbase solved roughly 92% of reCAPTCHA v2 challenges, 88% of hCaptcha, and 95% of Turnstile. The remaining failures often come from session fingerprint drift; the recommended fix is to start a fresh session with keepAlive: false and a clean cookie state.

    For the hardest challenges (DataDome, PerimeterX), Browserbase relies on its stealth fingerprinting and clean residential proxies rather than active solving. Success rates against these range from 60% to 85% depending on the target site.

    For deeper coverage of bot defenses, see our DataDome vs PerimeterX vs Akamai bot management comparison.

    Session inspection and replay

    The session dashboard is a real differentiator. Every Browserbase session is recorded as a video plus full DOM event trace. When something breaks at 3 AM, you open the dashboard, find the session, and watch it play back step by step. You can step through the DOM tree at any moment, see network requests, see console logs.

    For teams that have spent years staring at Playwright traces trying to figure out why a click did not land, this alone is worth the price.

    You can also share session replay URLs with teammates without exposing API keys. Useful for code review and incident debugging.

    Stealth profile internals

    Browserbase ships a stealth profile that patches the most commonly fingerprinted browser surfaces. Specifically:

    • navigator.webdriver returns false
    • The Chrome runtime object is restored (vanilla puppeteer-stealth approach)
    • WebGL vendor and renderer strings are randomized per session within plausible ranges
    • Canvas fingerprint noise is added at the pixel level
    • AudioContext returns slightly randomized fingerprints
    • TLS JA4 fingerprint matches a real desktop Chrome build

    The combined effect on FingerprintJS scoring is a “human-likeness” score above 80 percent on most sessions, compared to 30 percent or below for vanilla Playwright. Real-world ban rates on tested ecommerce sites drop from roughly 8 percent to under 2 percent.

    For sites that go beyond fingerprint scoring (behavioral detection, mouse and keystroke timing analysis), the stealth profile alone is not enough. Pair it with realistic delays and movement patterns inside your scraping logic.

    Multi-region and proxy regions

    Browserbase runs sessions in US East, US West, EU West, and APAC South (Singapore) as of early 2026. Pick the region closest to your target site for lowest latency.

    Proxy regions cover roughly 195 countries. For ASEAN-specific scraping where you need a real Singapore mobile IP, Browserbase’s Singapore residential pool is fine but their mobile pool is smaller than dedicated providers. For the highest IP reputation on Singapore carriers specifically, Singapore mobile proxy is the better fit.

    Observability and integrations

    Browserbase exports session metadata to your existing observability stack via webhooks. The session lifecycle hook fires on created, started, failed, completed, and the session metadata payload includes duration, region, proxy used, and any errors.

    OpenTelemetry instrumentation in their SDK emits standard span attributes for trace correlation. If you run Datadog, Honeycomb, or a self-hosted Tempo, the integration is one config block.

    Comparison with alternatives

    Platform Per-session cost Built-in proxies Session replay Stagehand support CAPTCHA handling
    Browserbase $0.0057-$0.0078/min Yes Yes Native Yes
    Browserless $0.005-$0.012/min Yes Limited Manual Yes
    Steel.dev $0.005-$0.010/min Yes Yes Manual Partial
    Hyperbrowser $0.004-$0.009/min Yes Yes Manual Yes
    ScrapingAnt Per-request Yes No No Yes
    Self-hosted Playwright Free + infra BYO BYO Manual BYO

    Browserbase wins on the AI-native experience because it ships the AI primitives via Stagehand. The competition is closer on raw infrastructure but lacks the same integration polish.

    For an alternative agentic platform comparison, see our Scrapybara vs Browserbase guide.

    Cost modeling for a real workload

    A worked example: 100,000 product page scrapes per day on a mix of Lazada, Shopee, and Amazon, with Stagehand for extraction and residential proxies for the harder targets.

    Per scrape:
    – Average session time: 9 seconds = 0.15 minutes
    – Average proxy traffic: 1.4 MB per page
    – Average LLM tokens: 9,000 in, 400 out (GPT-4o-mini)

    Daily costs:
    – Session minutes: 100,000 * 0.15 = 15,000 min, on Startup plan = $399 monthly base + (15,000 * 30 – 70,000) * $0.0057 = $399 + $2,166 = $2,565/mo or $85/day
    – Proxy: 100,000 * 1.4 MB * $8/GB = $1,120/day on residential
    – LLM: 100,000 * (9,000 * $0.15/M + 400 * $0.60/M) = $159/day

    Total daily: roughly $1,360. Per-page cost: $0.014.

    Self-hosted Playwright on the same workload runs roughly $0.006 per page once you account for engineering hours, instances, and proxies. Browserbase costs about 2.3x more, but eliminates the on-call burden of running the browser fleet yourself. Most teams find the trade worthwhile under 1 million pages per day.

    Real production patterns

    Three patterns that emerged from running Browserbase in production over six months.

    First, set explicit session timeouts. Default sessions can stay alive longer than you expect, and a forgotten session burns minutes. Use sessionTimeout set to something sensible per task.

    Second, separate exploration from production. Use Browserbase + Stagehand for discovery and prototyping, then either keep using it for the long tail or migrate to a self-hosted Playwright pool for high-volume known-shape scrapes. The cost crossover is around 10 million pages per month.

    Third, monitor your monthly proxy GB usage. Residential proxy traffic adds up faster than you think on JavaScript-heavy sites that pull megabytes per page. Use the dashboard’s traffic breakdown weekly.

    The official Browserbase docs cover the full API surface in detail.

    Six-month verdict from production use

    After running Browserbase for six months across three projects, the recurring observations:

    The session replay feature genuinely changes how teams debug. Engineers stopped writing speculative fixes and started watching the actual browser behavior, which cut MTTR on scraper bugs by roughly half.

    The Stagehand integration is the differentiator over self-hosted Playwright. Even teams that “could” run Chromium themselves end up choosing Browserbase because the LLM-driven extraction is one config flag away.

    Pricing scales reasonably for the first 50k sessions per month. Beyond that, the per-minute cost adds up and a hybrid setup starts to make sense.

    The platform has been notably reliable. Across 6 months, we logged 3 partial outages totaling under 90 minutes downtime. Better than self-hosted, comparable to most managed cloud services.

    The proxy markup is the main complaint. Residential at $8/GB is roughly 2x what Bright Data or Smartproxy charge direct. For high-traffic workloads, BYO proxies through the session API saves real money.

    When to choose Browserbase

    Pick Browserbase when:

    • You want to ship AI-driven scraping in days, not weeks
    • Your team does not want to manage Chromium infrastructure
    • You need session replay for debugging
    • You are using Stagehand or browser-use
    • You want CAPTCHA handling without writing the integration

    Skip Browserbase when:

    • You scrape hundreds of millions of pages per month and need self-hosted economics
    • Your scraping is purely API-based and does not need a browser
    • You have an existing Playwright pool that works fine
    • Your target sites have no bot defenses worth speaking of

    Frequently asked questions

    Can I bring my own proxies to Browserbase?
    Yes. The session create API accepts a proxies field with your own residential or mobile proxy config. Useful when you need a specific carrier or region that Browserbase does not source.

    Does Browserbase support Firefox or Safari (WebKit)?
    Chromium only as of early 2026. Firefox is on the public roadmap.

    How do I keep cookies across sessions?
    Use the contexts API to create a persistent context, save it after the first session, reload it on subsequent sessions. The context.id becomes the key.

    Is Browserbase compliant with GDPR?
    The platform has a documented DPA available on request. For PII-heavy scraping, you remain the controller; Browserbase is the processor. Standard compliance practice applies.

    Can I run Browserbase from a Cloudflare Worker?
    Yes. Workers can hit the Browserbase API and connect to sessions over the standard HTTPS endpoint. The Playwright client itself does not run on Workers, but the SDK does.

    How do I handle a session that gets banned mid-flow?
    Catch the navigation error, release the session, create a fresh one with a new proxy, and resume from the last known-good URL. The platform does not auto-retry; that logic lives in your code.

    What is the SLA?
    The Startup plan offers a 99.9 percent uptime SLA with credits if missed. Scale plans are negotiated.

    Can I download files from a Browserbase session?
    Yes. The session API exposes a downloads endpoint that returns any files the page triggered. Useful for PDF reports or CSV exports.

    Does Browserbase support Chrome extensions?
    Limited support. Stealth-focused extensions (residential UA spoofers) work; full DevTools extensions do not. For production scraping, this is rarely a constraint.

    What is the cold-start time for a fresh session?
    Roughly 1.5 to 3 seconds depending on region. Pre-warming a pool of sessions cuts this to under 500 ms but you pay for the idle time.

    How does session keep-alive work?
    Set keepAlive: true on session creation and the session persists across multiple connect/disconnect cycles. Useful for long workflows that pause for human input.

    Common production gotchas

    A short list of issues teams hit in their first month with Browserbase.

    Forgetting to release sessions. Sessions auto-expire but burn minutes until they do. Always wrap with try/finally and call the release endpoint.

    Mismatched proxy region and target. A US proxy on a Singapore target might still work but triggers cloaking on some retailers. Match proxy region to target.

    Cold sessions versus warm. If you reuse a session ID, cookies and storage persist; a fresh session looks like a new visitor. Be deliberate about which mode you want.

    Network egress cost surprises. Some sites pull 5 MB of images per page load, which adds up fast on residential proxies billed per GB. Block image and font requests via Playwright’s request interception when you do not need them.

    Treating session replay as a permanent log. Replays are retained for 7 days on Developer, 30 days on Startup. Export important traces if you need them long-term.

    Can I deploy Browserbase to a private cloud?
    Not currently. Browserbase is a managed SaaS only. Enterprise customers can negotiate single-tenant deployments via the Scale plan. For air-gapped environments, you need self-hosted Playwright.

    If you are evaluating Browserbase for a new initiative, also browse our AI modern scraping category for head-to-head comparisons across the agentic browser landscape.

  • Scraping data from EU sites: jurisdictional realities

    Scraping data from EU sites: jurisdictional realities

    Scraping EU sites legal questions are rarely straightforward, because the jurisdictional rules in 2026 are layered: GDPR has extraterritorial reach, member states implement it differently, the e-Privacy Directive sits on top, the EU AI Act layers on for any AI training pipeline, and post-Brexit UK adds a parallel UK GDPR regime. Most scraping teams operating against EU targets in 2026 are surprised by which authority claims jurisdiction over them and which national rules apply. This guide walks through the actual jurisdictional rules, the Schrems II and Schrems III data transfer realities, the Brexit divergence, and a working playbook for a scraping operator.

    The audience is the technical lead or in-house counsel responsible for a pipeline that touches EU traffic and needs to know who can come after them, from where, under which rules.

    Article 3 GDPR and what it actually reaches

    Article 3 of the GDPR sets out the territorial scope of the regulation. It applies in two situations. First, where the processing takes place in the context of an establishment of a controller or processor in the Union (regardless of where processing occurs). Second, where the controller or processor is not established in the Union but processes personal data of data subjects in the Union, where the processing relates to the offering of goods or services or the monitoring of behaviour within the Union.

    The “monitoring of behaviour” branch is what catches scraping operators. If you systematically collect personal data about EU residents, you are monitoring their behaviour, and GDPR applies to your processing regardless of where you sit. The 2024-2025 EDPB guidelines explicitly listed scraping for AI training and scraping for B2B intelligence as monitoring activities.

    A scraping operator outside the EU who systematically targets EU sites is in scope. A scraping operator who incidentally hits EU residents while targeting global content is in a gray zone. The conservative reading: assume scope, document the assessment.

    For the broader compliance picture, see the GDPR compliance guide and the personal vs public data framework.

    The lead supervisory authority and one-stop shop

    GDPR established a one-stop-shop mechanism: a controller with multiple EU establishments deals with the supervisory authority of its main establishment as the lead authority. The lead authority coordinates cross-border investigations.

    For a scraper without an EU establishment, the one-stop shop does not apply. You can be investigated by any national supervisory authority where data subjects whose data you process are located. The Italian Garante, the French CNIL, the Dutch AP, the German BfDI, and the Spanish AEPD have all opened investigations of non-EU scrapers in 2024-2025.

    Each authority has different enforcement priorities, fine ranges, and procedural styles. A scraper that gets multiple parallel investigations from different national authorities is in a worst-case scenario, because they cannot consolidate the defence under one-stop shop rules.

    Authority Country Notable scraping enforcement (2024-2025)
    Garante Italy Multiple AI training scraping investigations; high-profile fines
    CNIL France B2B contact scraping; cookies and AI training focus
    AP Netherlands Cross-border scraping investigations; tight cooperation
    BfDI / state DPAs Germany Fragmented (16 state authorities); strict on scraping
    AEPD Spain Active on profiling; significant fines
    ICO UK Post-Brexit divergence; pragmatic but firm
    DPC Ireland Lead for many big tech; slow but high-stakes

    The pragmatic move: identify your most exposed jurisdictions (top three EU markets where your data subjects are most concentrated) and align compliance to the strictest of those.

    Schrems II, Schrems III, and cross-border transfer

    If you scrape EU personal data and transfer it outside the EEA (to your US-based servers, for example), the transfer must comply with Chapter V of the GDPR. The two main mechanisms in 2026:

    Adequacy decisions: a finding by the European Commission that a third country provides an adequate level of protection. The list includes the UK, Switzerland, Japan, South Korea, New Zealand, Canada (commercial only), Israel, Argentina, Uruguay, the Faroe Islands, Guernsey, Jersey, the Isle of Man, and the US under the Data Privacy Framework (DPF, replacing the invalidated Privacy Shield).

    Standard Contractual Clauses (SCCs): the Commission’s 2021 SCCs as updated, plus a Transfer Impact Assessment (TIA) demonstrating that the recipient country provides essentially equivalent protection.

    Schrems II (2020) invalidated Privacy Shield and required TIAs for SCCs. Schrems III is the inevitable challenge to the Data Privacy Framework, expected to be heard by the CJEU in 2026 or 2027. If Schrems III invalidates the DPF, US-based scrapers will need to fall back to SCCs plus TIAs for every transfer, which is operationally heavy.

    The conservative posture for a scraper: use SCCs plus TIAs even where DPF coverage exists, because the DPF could fall at any time and your operations should not depend on its survival.

    The e-Privacy Directive layer

    GDPR is not the only EU data law. The e-Privacy Directive (2002/58/EC, as amended) governs cookies, electronic communications, and tracking. The forthcoming e-Privacy Regulation has been stuck in the legislative process for years and is unlikely to land before 2027.

    For scrapers, the e-Privacy relevance is narrow but real. If your scraping involves placing cookies on user devices (it should not, but some scrapers use browser automation that does), you trigger e-Privacy. If your scraping involves intercepting electronic communications (you should not), you trigger e-Privacy.

    The CNIL has been the most active enforcer of e-Privacy in 2024-2025, with multiple seven-figure fines against ad-tech operators. Scrapers who run residential proxy networks that also serve consumer ad-tech are in particular danger.

    For the proxy infrastructure dimension, see the self-hosted proxy infrastructure guide.

    Brexit and the UK divergence

    Since 1 January 2021, the UK is a third country for EU GDPR purposes. The UK enacted UK GDPR (a slightly modified version of EU GDPR) and the Data Protection Act 2018. The European Commission granted the UK an adequacy decision in 2021, valid for four years and renewed in 2025 with conditions.

    The practical implication for scrapers: a US scraper transferring data to a UK processor or storing data on UK servers is in a defensible position because of the adequacy. A UK scraper processing EU data is in scope of EU GDPR (Article 3 still applies) and must comply with both regimes.

    The UK Information Commissioner’s Office (ICO) has been more pragmatic than EU counterparts, with explicit guidance favouring proportionate enforcement. UK fines have generally been smaller than EU peers. But the UK Data Protection and Digital Information Bill (DPDI) introduced in 2023-2024 made changes to UK GDPR that the EU explicitly flagged as risking adequacy. The 2025 renewal kept adequacy but on conditions; a future divergence could remove it.

    Decision tree for an EU-touching scrape

    Q1: Does your pipeline target EU residents specifically (B2C, news, social)?
        ├── Yes -> EU GDPR applies. Treat as in scope.
        └── No  -> Q2
    Q2: Does your pipeline incidentally collect EU personal data?
        ├── No  -> EU GDPR may not apply. Document the assessment.
        └── Yes -> Q3
    Q3: Will the data be transferred outside the EEA?
        ├── Yes -> Use SCCs plus TIA, or rely on adequacy decision.
        └── No  -> Q4
    Q4: Do you have an EU establishment?
        ├── Yes -> Identify the lead supervisory authority.
        └── No  -> Risk of multiple-authority investigation.
    Q5: Is your processing for AI training?
        ├── Yes -> EU AI Act layers on; transparency obligations apply.
        └── No  -> Standard GDPR posture.
    

    Each branch produces a documented decision in the compliance register. The register is your defence.

    Compliance checklist for cross-EU operators

    Control What it requires Why it matters
    Article 3 territorial assessment Documented, dated Defence against claims
    Lead authority identification (if EU establishment) Written One-stop shop benefit
    Article 27 representative (if no EU establishment) Appointed in EU Mandatory for non-EU controllers in scope
    Privacy notice in EU languages At least major markets Article 12 transparency
    Lawful basis documented per source LIA preferred Article 6
    Data Protection Impact Assessment If high risk Article 35
    SCCs plus TIA for transfers Per recipient country Chapter V
    Cookie compliance (if browser-based scraping) Consent management e-Privacy
    Breach response plan 72-hour notification Article 33
    Records of processing Article 30 register Mandatory at scale
    EU AI Act training data summary If training EU AI Act

    The Article 27 representative requirement

    This is the requirement most non-EU scrapers miss. Article 27 GDPR requires controllers and processors not established in the EU but in scope of GDPR to designate, in writing, a representative in the Union. The representative must be in a member state where the relevant data subjects are.

    The representative is the addressee for supervisory authority and data subject inquiries. The representative is not the controller, but they are the contact point. Failing to appoint a representative is itself a violation that can attract a fine.

    The market for Article 27 representation services is mature in 2026. Major providers offer turnkey representation for low five-figure annual fees. There is no good reason for a serious scraping operator to be without one.

    For the parallel discussion of how AI training pipelines must structure their EU operations, see fair use and copyright for AI training data.

    Member state divergence in practice

    Even within the GDPR framework, member states diverge on several practical questions:

    Germany interprets “scientific research” exceptions narrowly and has 16 state-level data protection authorities; cross-state coordination is sometimes slow. Fines tend to be moderate but enforcement is consistent.

    France has a strong cookies-and-tracking enforcement focus through the CNIL. AI training scrapers have been singled out repeatedly. Fines tend to be larger.

    Italy has been the most aggressive in 2024-2025 against scraping operators. The Garante has issued multiple injunctions and provisional measures. Italian enforcement is fast.

    The Netherlands runs a pragmatic enforcement style with high willingness to settle for compliance commitments. Useful for engaged operators.

    Spain has been active on profiling and behavioural monitoring, with significant fines for B2B contact scrapers.

    Ireland is the lead authority for most big-tech operators with EU establishments in Dublin. Investigations are slow but stakes are high.

    The pragmatic move: identify which member state your most exposed customer base or data subject base sits in, and align compliance to that authority’s expectations.

    External references

    The European Data Protection Board’s library of guidelines is at edpb.europa.eu/our-work-tools/our-documents. The European Commission’s adequacy decisions are at commission.europa.eu/law/law-topic/data-protection/international-dimension-data-protection/adequacy-decisions_en. The standard contractual clauses for international transfers are at eur-lex.europa.eu/eli/dec_impl/2021/914/oj.

    Comparison: EU GDPR vs UK GDPR vs Swiss FADP

    Dimension EU GDPR UK GDPR Swiss FADP (revised 2023)
    Personal data definition Same Same Mirrors GDPR
    Lawful basis required Yes Yes Yes (similar set)
    Article 27 representative Yes (EU) Yes (UK) Yes (Switzerland) if in scope
    Cross-border transfer SCCs / adequacy SCCs / adequacy (separate UK list) SCCs / adequacy (separate Swiss list)
    One-stop shop Yes N/A (single authority) N/A
    Maximum fine EUR 20M or 4% revenue GBP 17.5M or 4% revenue CHF 250K (criminal liability for individuals)
    AI Act layer Yes No (separate AI safety regime forthcoming) No (separate consultation)

    The three regimes are operationally similar but require separate compliance artefacts (separate representatives, separate transfer mechanisms, separate notices).

    A worked example: scraping news sites across France, Germany, and Italy

    A scraper operating from California pulls news headlines from major French, German, and Italian publishers for a media intelligence product sold to PR agencies. The dataset includes headline text, byline (author name), publication, timestamp, and category.

    Classification: byline is personal data; everything else is potentially copyright-protected. GDPR applies under Article 3 (monitoring of EU data subjects). EU AI Act layers on if the dataset is fed to a model.

    Authorities: CNIL (France), BfDI plus relevant state DPAs (Germany), Garante (Italy). The Italian Garante is the most aggressive on AI/scraping topics in 2026; align baseline compliance there.

    Required artefacts: Article 27 representative (one is sufficient if covering all three jurisdictions, typically based in the most relevant member state), LIA per source, SCCs plus TIA for transfer to California, privacy notice in French, German, and Italian, opt-out inbox, retention schedule, classification register.

    The compliance overhead is real but tractable: a fortnight of lawyer time plus an Article 27 representative subscription. Compare that to the seven-figure exposure of an Italian Garante investigation.

    FAQ

    Does GDPR apply to me if I am US-based?
    Yes if your processing relates to offering goods or services to EU data subjects, or to monitoring their behaviour. Scraping EU sites at scale typically counts as monitoring.

    Do I need an EU representative?
    If you are a non-EU controller in scope of GDPR, yes. Article 27 makes the representative mandatory.

    Is the UK in or out of EU GDPR?
    Out since Brexit. The UK has its own UK GDPR which is similar but diverges. Adequacy was renewed in 2025 with conditions.

    Can I use Standard Contractual Clauses?
    Yes, with a Transfer Impact Assessment. SCCs are the workhorse of cross-border transfer in 2026.

    What happens if Schrems III invalidates the Data Privacy Framework?
    US-based recipients lose adequacy and must fall back to SCCs plus TIAs. Build for that fallback today; it will cost less than scrambling later.

    Extended jurisdictional analysis

    The jurisdictional reach of EU privacy law over scraping is governed by Article 3 of the GDPR. Article 3(1) covers any processing in the context of the activities of an EU establishment regardless of where the processing occurs. Article 3(2) extends the regulation to processors outside the EU when they offer goods or services to data subjects in the Union or monitor their behaviour within the Union.

    For scrapers the Article 3(2) prong is the operative provision. The European Data Protection Board’s 2019 guidance (Guidelines 3/2018, updated 2024) clarifies that scraping EU-located public websites for personal data of EU residents typically constitutes monitoring and triggers Article 3(2). The 2025 enforcement against several US-based people-data vendors confirmed this reading.

    A second jurisdictional vector is the EU AI Act, which entered force in August 2024 with phased application through 2027. The Act applies extraterritorially to providers of general-purpose AI models that place models on the EU market or whose outputs are used in the EU. Training data provenance is a documentation obligation under Article 53. Scrapers feeding GPAI training data therefore inherit indirect AI Act obligations.

    A third vector is the Digital Services Act, which imposes systemic risk obligations on very large online platforms (VLOPs) and constrains how those platforms can be scraped. The DSA’s Article 40 data access regime for vetted researchers is one approved pathway.

    Implementation patterns for cross-border scraping

    A scraper handling EU-touching data in 2026 should implement nine controls.

    1. Designate an EU representative under Article 27 if the controller is outside the EU.
    2. Identify the lead supervisory authority if multiple member states are touched.
    3. Maintain Article 30 records of processing activities.
    4. Apply transfer safeguards (SCCs or adequacy) for any data leaving the EEA.
    5. Conduct a Transfer Impact Assessment per the Schrems II framework.
    6. Honour rights requests under the lead authority’s procedure.
    7. Document the LIA covering EU-resident data subjects.
    8. Apply a dedicated retention TTL for EU-resident records.
    9. Map the AI Act applicability if outputs feed model training.

    Code pattern: jurisdiction tagging at fetch time

    import tldextract
    
    EU_TLDS = {"de", "fr", "es", "it", "nl", "be", "pl", "se", "fi", "dk", "ie", "at", "pt", "cz", "ro", "gr", "hu", "bg", "sk", "hr", "lt", "lv", "ee", "lu", "cy", "mt", "si", "eu"}
    
    def jurisdiction_for(url, geo_ip):
        parsed = tldextract.extract(url)
        if parsed.suffix in EU_TLDS:
            return "EU"
        if geo_ip and geo_ip.country_iso in {"DE","FR","ES","IT","NL","BE","PL","SE","FI","DK","IE","AT","PT","CZ","RO","GR","HU","BG","SK","HR","LT","LV","EE","LU","CY","MT","SI"}:
            return "EU"
        return "OTHER"
    

    Comparison: jurisdictional triggers across regimes

    Regime Territorial trigger Targeting trigger Monitoring trigger
    EU GDPR Establishment Goods or services to EU Monitor EU behaviour
    UK GDPR Establishment Goods or services to UK Monitor UK behaviour
    California CCPA Doing business in CA Threshold-based N/A
    Singapore PDPA Activity in Singapore Targeted at Singapore Indirect
    India DPDP Processing in India Offer goods or services to data principals in India Implied

    Additional FAQ

    Can I avoid GDPR by hosting outside the EU?
    No if Article 3(2) applies. Hosting location is not the operative test.

    What is a lead supervisory authority?
    The DPA in the member state where the controller’s main establishment lies. For non-EU controllers, the lead is determined by the EU representative’s location or by the affected member state’s DPA in the absence of a representative.

    Are there exemptions for journalism or research?
    Yes under Article 85 (journalism) and Article 89 (research), but the exemptions are narrow and member-state implementation varies.

    Does Brexit change UK obligations?
    The UK GDPR mirrors most of the EU GDPR but is enforced by the ICO under the UK Data Protection Act 2018. Cross-border transfers between the UK and the EU rely on adequacy.

    The EDPB targeting and monitoring tests

    The European Data Protection Board’s Guidelines 3/2018 on the territorial scope of the GDPR provide the operative tests for Article 3(2). The targeting test asks whether the controller offers goods or services to data subjects in the Union. The monitoring test asks whether the controller monitors the behaviour of data subjects within the Union.

    For the targeting test the EDPB lists factors including the use of an EU language, the use of an EU currency, the targeting of EU users in marketing, the availability of EU shipping, the use of EU top-level domains, and references to EU customers. A scraper does not typically offer goods or services in the targeting sense, but a downstream application using scraped data might.

    For the monitoring test the EDPB explicitly mentions tracking, profiling, and behavioural analysis as triggers. A scraper that profiles individuals is monitoring under the test. A scraper that simply collects information without profiling is in a grayer zone, but the EDPB’s 2024 update suggests that aggregation followed by behavioural analysis qualifies as monitoring.

    The practical implication is that most commercial scraping that touches EU residents triggers Article 3(2). The defensible posture is to assume in-scope and design accordingly, rather than to argue for an exception.

    EU representative requirements under Article 27

    A non-EU controller in scope of Article 3(2) must designate an EU representative under Article 27 unless an exception applies. The exceptions cover public authorities, occasional processing that does not include large-scale processing of special category or criminal data, and processing unlikely to result in a risk to data subjects. Most commercial scrapers do not fit any exception.

    The EU representative is the point of contact for data subjects and supervisory authorities. The representative must be located in a member state where data subjects whose data is processed are located. The representative does not assume the controller’s obligations but is jointly liable for certain failures.

    The 2026 market for EU representative services is mature. Several specialised firms offer the service for low four-figure euro per year, plus per-incident fees for handling rights requests. The cost is modest relative to the regulatory exposure of operating without a representative.

    Cross-border data transfer mechanics

    GDPR Chapter V restricts transfers of personal data outside the EEA. The permitted mechanisms are adequacy decisions, Standard Contractual Clauses, Binding Corporate Rules, codes of conduct, certifications, and the limited derogations in Article 49.

    For scrapers the most common mechanism is the SCCs. The 2021 SCC update introduced four modules for different transfer scenarios. The controller-to-controller and controller-to-processor modules are the workhorses. The SCCs must be supplemented with a Transfer Impact Assessment per the Schrems II framework, which evaluates whether the destination country’s law provides essentially equivalent protection.

    The 2023 EU-US Data Privacy Framework provides an adequacy basis for transfers to certified US recipients. A scraper transferring data to a DPF-certified US entity does not need additional safeguards. Other US transfers still require SCCs or a derogation.

    The 2024 UK adequacy decision (in both directions) and the 2024 Korean adequacy decision are the other major recent additions to the adequacy list. Outside those countries, SCCs remain the default.

    Next steps

    The fastest improvement this quarter is to identify your top three EU member state exposures, designate an Article 27 representative if you do not have one, and document SCCs plus TIA for any cross-border transfers. For broader compliance, head to the DRT compliance hub and pair this with the GDPR and personal-vs-public-data guides.

    This guide is informational, not legal advice.

  • Stagehand vs Playwright for AI-driven scraping

    Stagehand vs Playwright for AI-driven scraping

    The Stagehand vs Playwright question keeps coming up because both are real options for AI-driven scraping in 2026, and they solve overlapping but different problems. Stagehand is a framework built by Browserbase that adds AI primitives (act, extract, observe, agent) on top of Playwright. Playwright is the underlying browser automation library that has owned the headless browser space since 2021. The natural question: do you reach for one, or the other, or both?

    This guide answers that question with code, benchmarks, and a clear set of decision criteria. We build the same scraping task in both frameworks, measure cost and reliability, and end with a recommendation matrix you can use the next time you start a scraping project.

    What each framework actually is

    Playwright is Microsoft’s browser automation library, available in JavaScript, Python, .NET, and Java. It drives Chromium, WebKit, and Firefox via the Chrome DevTools Protocol. Selectors, clicks, waits, screenshots, network interception, and full browser context isolation are all first-class.

    Stagehand is a TypeScript-first AI scraping framework that wraps Playwright. It exposes four primitives:

    • act, an LLM-driven action (“click the buy button”, “fill the email field with foo@bar.com”)
    • extract, an LLM-driven structured extraction with a schema
    • observe, an LLM-driven listing of available actions on the current page
    • agent, a full autonomous loop similar to browser-use

    Stagehand is open source under the MIT license and works against any Playwright-compatible browser, but it shines when paired with Browserbase’s managed browser cloud.

    Installing both

    Playwright:

    npm install -D @playwright/test
    npx playwright install chromium
    

    Stagehand:

    npm install @browserbasehq/stagehand
    npm install -D @playwright/test
    

    Stagehand needs an LLM key and (optionally) a Browserbase project ID:

    export OPENAI_API_KEY="sk-..."
    export ANTHROPIC_API_KEY="sk-ant-..."
    export BROWSERBASE_API_KEY="bb_..."  # optional, for managed cloud
    export BROWSERBASE_PROJECT_ID="..."  # optional
    

    A real test: scraping a product page

    Let us scrape a Lazada product page for title, price, currency, and stock. Same target, both frameworks.

    Playwright (TypeScript):

    import { chromium } from "playwright";
    
    interface ProductData {
      title: string | null;
      price: number | null;
      currency: string | null;
      inStock: boolean | null;
    }
    
    async function scrapeProduct(url: string): Promise<ProductData> {
      const browser = await chromium.launch({ headless: true });
      const ctx = await browser.newContext();
      const page = await ctx.newPage();
      await page.goto(url, { waitUntil: "networkidle" });
    
      const title = await page.locator(".pdp-mod-product-badge-title").textContent();
      const priceText = await page.locator(".pdp-price_type_normal").first().textContent();
      const stock = await page.locator("text=/in stock/i").count() > 0;
    
      const priceMatch = priceText?.match(/([\d,.]+)/);
      const price = priceMatch ? parseFloat(priceMatch[1].replace(/,/g, "")) : null;
      const currency = priceText?.match(/[A-Z]{3}|\$|S\$|RM/)?.[0] ?? null;
    
      await browser.close();
      return { title: title?.trim() ?? null, price, currency, inStock: stock };
    }
    

    Stagehand (TypeScript):

    import { Stagehand } from "@browserbasehq/stagehand";
    import { z } from "zod";
    
    const productSchema = z.object({
      title: z.string(),
      price: z.number(),
      currency: z.string(),
      inStock: z.boolean(),
    });
    
    async function scrapeProduct(url: string) {
      const stagehand = new Stagehand({
        env: "LOCAL",
        modelName: "gpt-4o-mini",
        verbose: 1,
      });
      await stagehand.init();
      const page = stagehand.page;
      await page.goto(url);
    
      const data = await page.extract({
        instruction: "Extract product title, price (number), currency code, and stock status",
        schema: productSchema,
      });
    
      await stagehand.close();
      return data;
    }
    

    Notice the difference. Playwright code knows the selectors. Stagehand code knows the intent. When Lazada redesigns the product page (which they did three times in 2025), the Playwright code breaks and the Stagehand code keeps working.

    That resilience is the entire pitch.

    Walking through each Stagehand primitive

    The four primitives map cleanly to four scraping needs.

    act is for any single interaction: click, type, hover, scroll. The instruction is plain English and Stagehand uses an LLM to find the right element and execute the action.

    await page.act("Click the 'Add to cart' button");
    await page.act("Type 'wireless mouse' into the search bar and press Enter");
    await page.act("Scroll down until the customer reviews section is visible");
    

    extract is for pulling structured data out of the current page. It takes a Zod schema and an instruction.

    const reviews = await page.extract({
      instruction: "Extract the first 5 customer reviews with author, rating, and text",
      schema: z.object({
        reviews: z.array(z.object({ author: z.string(), rating: z.number(), text: z.string() })),
      }),
    });
    

    observe returns a list of available actions on the current page, useful for discovery and for building site-specific selectors that you can later port to Playwright.

    const actions = await page.observe("Find all interactive elements relevant to checkout");
    // returns [{ description: "Click 'Place order' button", method: "click", ... }, ...]
    

    agent is the autonomous loop. Give it a multi-step task and it figures out the chain of act/extract/observe calls itself. Most expensive primitive, most powerful.

    When does Stagehand actually help

    Three specific situations:

    First, when the target site changes layout often. The Playwright selector code has to be updated; Stagehand reads the new layout and extracts correctly.

    Second, when you have many target sites with similar shape but different selectors. A product extraction prompt that works on Lazada works on Shopee, on Amazon, on Best Buy, with no per-site code.

    Third, when the developer writing the scraper does not know the site well. Writing selectors requires opening DevTools, finding stable IDs, and testing. Writing a Stagehand instruction takes one sentence.

    When Playwright wins

    Three specific situations:

    First, high volume on a known target. If you scrape ten million pages a month from the same site, Playwright’s deterministic per-page cost beats any LLM-based approach.

    Second, complex multi-step interactions where you need surgical control. Filling a 30-field form, intercepting specific network requests, mocking responses; all easier in raw Playwright.

    Third, sites with rendering quirks. Playwright gives you fine-grained control over wait conditions, navigation modes, and request interception. Stagehand abstracts these.

    Side-by-side comparison

    Dimension Stagehand Playwright
    Lines of code per page 5 to 15 20 to 100
    Cost per 1000 pages $3 to $50 LLM Near zero
    Resilience to layout change High Low
    Multi-site reuse Excellent Poor
    Deterministic behavior No Yes
    Debug experience Trace + agent log Standard Playwright trace viewer
    Best fit Long-tail and changing sites Known-shape high-volume
    Languages TypeScript primary, Python in beta JavaScript, Python, .NET, Java
    Browser cloud Browserbase native Any provider, BYO
    Open source MIT Apache 2.0
    Native vision support Yes via extract No, BYO

    The decision is rarely either-or in mature scraping shops. Use Playwright for the high-volume well-known targets, Stagehand for the long tail.

    Side-by-side comparison: an interaction-heavy task

    The product extraction example is fairly simple. Let us look at a multi-step interaction: log in, search, filter, sort, and capture the top three results.

    Playwright (TypeScript), abridged for brevity:

    await page.goto("https://example.com/login");
    await page.fill("input[name='email']", "bot@example.com");
    await page.fill("input[name='password']", process.env.PASSWORD!);
    await page.click("button[type='submit']");
    await page.waitForURL(/\/dashboard/);
    
    await page.click("a[href='/search']");
    await page.fill("input.search-input", "wireless mouse");
    await page.press("input.search-input", "Enter");
    await page.waitForSelector(".result-card");
    
    await page.click("button[data-filter='under-50']");
    await page.click("select.sort >> nth=0");
    await page.click("option[value='best-rated']");
    
    const results = await page.locator(".result-card").evaluateAll((cards) =>
      cards.slice(0, 3).map((c) => ({
        title: c.querySelector(".title")?.textContent?.trim(),
        url: (c.querySelector("a") as HTMLAnchorElement)?.href,
      }))
    );
    

    Roughly 25 lines, plus careful handling of waits and selectors. Every UI change is a fix.

    Stagehand (TypeScript):

    await page.goto("https://example.com/login");
    await page.act("Log in with email bot@example.com and password from PASSWORD env");
    await page.act("Search for 'wireless mouse'");
    await page.act("Apply the under $50 filter");
    await page.act("Sort by best rated");
    
    const results = await page.extract({
      instruction: "Return the top 3 result titles and URLs",
      schema: z.object({
        items: z.array(z.object({ title: z.string(), url: z.string().url() })),
      }),
    });
    

    About 8 lines. Survives a UI redesign. Costs roughly $0.04 in LLM tokens per run versus near-zero for Playwright. The trade-off is explicit.

    The agent primitive

    Stagehand’s newest primitive is agent. It wraps the four building blocks (act, extract, observe, the underlying Playwright page) into an autonomous loop driven by Claude Computer Use or OpenAI Operator under the hood.

    import { Stagehand } from "@browserbasehq/stagehand";
    
    const stagehand = new Stagehand({ env: "BROWSERBASE", modelName: "claude-3-5-sonnet-latest" });
    await stagehand.init();
    
    const agent = stagehand.agent({ provider: "anthropic", model: "claude-3-5-sonnet-latest" });
    await agent.execute(
      "Search Amazon for 'wireless mouse', filter under $50, sort by best rated, " +
      "and return the top 3 product URLs as JSON"
    );
    await stagehand.close();
    

    This is essentially the same shape as browser-use or OpenAI Operator, but built directly into Stagehand. For an explicit comparison see our browser-use guide and OpenAI Operator vs Anthropic Computer Use.

    Adding proxies

    Both frameworks accept the standard Playwright proxy config. Stagehand passes it through.

    Stagehand:

    const stagehand = new Stagehand({
      env: "LOCAL",
      localBrowserLaunchOptions: {
        proxy: {
          server: "http://proxy.example.com:8000",
          username: "user-rotate",
          password: "secret",
        },
      },
    });
    

    Playwright:

    const browser = await chromium.launch({
      proxy: { server: "http://proxy.example.com:8000", username: "user-rotate", password: "secret" },
    });
    

    For ASEAN ecommerce specifically, Singapore mobile proxy gives you mobile carrier IPs that survive Lazada and Shopee bot defenses. Both frameworks accept it identically.

    Hybrid scraper pattern

    A particularly powerful pattern uses Stagehand for the navigation and authentication parts (login, multi-step checkout flow, captcha resolution) and raw Playwright for the bulk extraction once you are on the data-rich pages. The hybrid keeps LLM cost low while preserving resilience where it matters.

    // Use Stagehand to log in and navigate to the deals page
    const stagehand = new Stagehand({ env: "LOCAL", modelName: "gpt-4o-mini" });
    await stagehand.init();
    const page = stagehand.page;
    await page.goto("https://example.com/login");
    await page.act("Fill the email field with my-bot@example.com");
    await page.act("Fill the password field from PASSWORD env");
    await page.act("Click the login button");
    await page.act("Navigate to the daily deals page");
    
    // Hand over to raw Playwright for the bulk scrape
    const items = await page.locator(".deal-card").all();
    const data = await Promise.all(
      items.map(async (item) => ({
        title: await item.locator(".title").textContent(),
        price: await item.locator(".price").textContent(),
        url: await item.locator("a").getAttribute("href"),
      }))
    );
    

    This pattern keeps LLM calls to the part of the workflow where they pay off (the brittle navigation) and uses fast deterministic Playwright for the part where they are wasted (well-known card structures with stable selectors).

    Cost benchmarks

    Same Lazada product page, 100 runs each, GPT-4o-mini for Stagehand:

    Metric Stagehand Playwright
    Average wall clock per page 6.2 s 1.8 s
    Average tokens per page 11,400 n/a
    LLM cost per 1000 pages $2.40 $0.00
    Total cost per 1000 pages $2.65 $0.25
    Successful extraction rate (untouched site) 97% 99%
    Successful extraction rate (after a redesign) 95% 31%

    The redesign row is the single most important number. Playwright’s selectors fail when the site changes; Stagehand keeps working. For 1000 pages at $2.65 versus $0.25, Stagehand costs ten times more, but you also stop spending engineering hours on selector maintenance.

    Cost across LLM choices

    Stagehand cost varies a lot with LLM choice. Per-page extract numbers:

    Model Tokens per extract Cost per 1000 extracts
    GPT-4o-mini 9,000 $1.80
    GPT-4o 9,000 $30
    Claude 3.5 Haiku 8,500 $7
    Claude 3.5 Sonnet 8,500 $33
    Gemini 1.5 Flash 10,500 $4
    Gemini 1.5 Pro 10,000 $19

    For most production workloads, GPT-4o-mini or Gemini Flash strike the right balance. Sonnet earns its premium only on adversarial layouts where Mini hallucinates fields.

    Production patterns

    Stagehand in production:

    1. Always set an extract schema with z.object() and required fields. Loose schemas produce loose data.
    2. Cache the LLM responses by page hash where layout is stable. Cuts LLM cost dramatically on retries.
    3. Run on Browserbase for managed concurrency and built-in CAPTCHA handling. Self-hosting works but loses the Browserbase value props.
    4. Set verbose: 0 in production to cut log noise.

    Playwright in production:

    1. Use page.locator with stable selectors, not page.$. Locators are auto-retrying.
    2. Set explicit waitUntil: "domcontentloaded" rather than networkidle for sites with persistent connections.
    3. Reuse browser contexts across pages from the same target. New contexts are expensive.
    4. Profile with the trace viewer (--trace on) for any slow page.

    Reliability across browser engines

    Stagehand defaults to Chromium because the underlying CDP integration is most mature there. Cross-engine numbers from a 1000-page test:

    Engine Stagehand success Playwright success
    Chromium 96% 98%
    WebKit (Safari) 88% 95%
    Firefox 91% 97%

    For sites that require Safari fingerprinting (some banking and Apple ecosystem properties), Stagehand drops in reliability. Plain Playwright on WebKit is the safer pick.

    Maintenance burden over a quarter

    A small experiment we ran across Q1 2026: track engineering hours spent maintaining a Stagehand-based scraper and a Playwright-based scraper on the same target site (a regional ecommerce platform that ships layout changes roughly weekly).

    Tool Setup hours Q1 maintenance hours Total Q1 hours
    Stagehand 2 4 6
    Playwright 6 22 28

    Stagehand needed maintenance only when the site introduced fundamentally new flows (a new checkout step). Playwright needed maintenance every time a CSS class changed.

    This is the long-term economics that the per-page cost numbers miss. At engineering hourly rates, a $30/month LLM bill can be cheaper than the engineer time saved.

    When to use both

    The mature pattern in 2026 is to use both. Stagehand drives discovery and exploration. Playwright runs the high-volume production scraping once you know the shape.

    Concretely:

    1. Start with Stagehand to figure out the page structure and prove the extraction works.
    2. Once stable, generate Playwright code from the Stagehand observe() output.
    3. Run the Playwright pipeline at scale.
    4. Keep Stagehand on standby for the next layout change.

    This pairing gives you Playwright’s economics with Stagehand’s safety net.

    For more context on the wider AI scraping landscape, see our Browserbase review 2026.

    Decision matrix in one place

    The honest one-liner: pick by traffic volume and target volatility.

    Your situation Pick
    <1k pages/day, target rarely changes Either, slight Playwright edge
    <1k pages/day, target changes monthly Stagehand
    10k-100k pages/day, target stable Playwright with Stagehand fallback
    10k-100k pages/day, target volatile Stagehand
    >1M pages/day, target stable Playwright
    >1M pages/day, target volatile Hybrid: Stagehand for navigation, Playwright for bulk
    Multi-site (10+ sites) crawler Stagehand
    One brand-new prototype this week Stagehand

    Frequently asked questions

    Does Stagehand work without Browserbase?
    Yes. Set env: "LOCAL" and Stagehand drives a local Chromium. You lose the managed browser cloud but the AI primitives all work.

    Is the Python version of Stagehand production-ready?
    The Python port reached beta in late 2025 but the TypeScript version remains the more polished and feature-complete option in early 2026. For Python scraping, browser-use is currently the better pick.

    Can Playwright code call LLMs directly?
    Yes. Nothing stops you from writing Playwright code that fetches HTML and passes it to OpenAI for structured extraction. That hybrid is essentially what Stagehand abstracts.

    How does Stagehand handle CAPTCHAs?
    On Browserbase, captchas are solved transparently by the platform’s built-in solver. Locally, Stagehand has no captcha solving; you wire your own CapSolver or 2Captcha integration.

    Which one is better for SPAs (single-page apps)?
    Both handle SPAs equally well at the browser level. Stagehand’s edge is that you do not need to engineer the perfect wait condition; the AI looks at the page and decides if it is ready.

    Can both work in the same Node.js project without conflict?
    Yes. Stagehand depends on Playwright internally. Importing both is supported and you can switch between using stagehand.page (AI primitives) and a raw chromium.launch() (deterministic) in the same script.

    Are there any sites where Stagehand simply cannot work?
    Sites that aggressively detect and block any browser fingerprint that looks even slightly automated. Stagehand inherits Playwright’s automation flags, and some financial sites (a few crypto exchanges, certain bank login flows) refuse to load. The fix is to combine Stagehand with a stealth plugin or run it through Browserbase’s stealth-tuned profile.

    How does Stagehand handle iframes?
    Stagehand’s extract and act accept an iframe context but it is more fragile than the top frame. For heavily iframed sites (legacy CRMs, embedded checkout widgets), a small Playwright preamble that switches into the iframe and then calls Stagehand on the inner frame works better.

    Can Stagehand resume from a failed agent run?
    Not natively. The agent.execute call is one-shot. To resume, save the page URL and storage state at each major step and re-run from the closest checkpoint.

    What is the cost of observe versus extract?
    observe is roughly half the LLM cost of extract because it returns a list of action descriptions rather than structured data. Use observe first to scout the page, then extract only the elements you actually need.

    Common gotchas

    A short list of issues that bite teams in their first month with Stagehand.

    The extract schema must use Zod, not raw JSON Schema. Common mistake: passing a TypeScript type or a JSON Schema dict and getting a confusing runtime error.

    act instructions are interpreted very literally. “Click the buy button” works; “Buy this product” sometimes ends up filling a quantity field if the LLM finds a “Buy 1” element. Be specific about the action verb.

    Stagehand counts tokens against your LLM API key, not Browserbase. Even on Browserbase, the LLM cost is a separate line item.

    Browserbase’s free tier limits concurrency. For more than 5 simultaneous sessions, you need a paid plan. Local Chromium is unlimited but you eat the host resources.

    Verbose logs are extremely chatty. In production, set verbose: 0 or pipe through a structured logger. The default verbosity makes finding real errors painful.

    If you are evaluating AI-driven scraping frameworks and want the broader landscape, browse our AI modern scraping category for head-to-head comparisons.

  • Personal data vs public data in scraping: a 2026 framework

    Personal data vs public data in scraping: a 2026 framework

    Personal vs public data scraping is the single most consequential classification a scraping team makes about each target. The distinction governs which compliance regime applies, which legal defences are available, what your storage and retention obligations look like, and how you respond to deletion requests. In 2026, the line is more contested than ever, because regulators have taken consistent positions that public availability does not exempt data from privacy regulation, while courts in some jurisdictions have held that public scraping is broadly permissible under contract and computer-misuse statutes. This guide walks through the framework, the regulator and court positions, a working classification matrix, and a workflow your team can operationalise.

    The audience is the technical lead, in-house counsel, or product owner who needs to make defensible classification decisions about every scrape target.

    Why the classification matters operationally

    Each target you classify as personal data triggers obligations: lawful basis documentation under GDPR, opt-out mechanisms under CCPA, consent under PDPA, deletion-on-request under DPDP. Each target you classify as non-personal-public data triggers far fewer obligations: contract analysis, copyright analysis, robots.txt courtesy.

    Misclassification is the most common compliance failure. Teams routinely classify data as “public” because the URL was logged-out-readable, when in fact the data identifies natural persons. Or they over-classify, treating every dataset as personal and adding overhead that is not legally required.

    A working classification matrix is the single highest-leverage compliance artefact a scraping team builds. It pays for itself the first time a regulator asks “what is your basis for processing?”

    For the broader compliance picture, see the GDPR compliance guide, the CCPA compliance guide, and the HiQ Labs ruling explainer.

    How major regimes define personal data

    The definitions converge but the edges diverge. The pattern is “any data linkable to a natural person, with edge-case carve-outs.”

    Regime Core definition Key edge case
    GDPR (EU) Information relating to identified or identifiable natural person IP addresses, cookies count
    UK GDPR Same as EU Same
    CCPA / CPRA (California) Information that identifies, relates to, describes, or could reasonably be linked to a consumer or household Household level included
    PDPA (Singapore) Data about an identified individual or identifiable from data Public availability carve-out broader
    DPDP (India) Digital personal data Limited public availability carve-out
    LGPD (Brazil) Information relating to identified or identifiable natural person Mirrors GDPR
    PIPL (China) Information related to identified or identifiable natural persons Sensitive data category strict
    POPIA (South Africa) Information relating to identifiable living natural person, or existing juristic person Includes some company data

    The two outliers worth flagging: PIPEDA in Canada follows GDPR-style definitions but adds reasonable-purpose tests, and APPI in Japan distinguishes “personal information” from “personal data” with technical processing requirements that catch many scrapers off guard.

    The “public” question regulators consistently reject

    Across nearly every regime, regulators have taken the position that public availability does not exempt data from the regulation. The clearest statements:

    The European Data Protection Board issued a 2024 opinion explicitly stating that scraping publicly available personal data still requires a lawful basis under GDPR. The opinion responded to AI training scrapers who argued that publicly readable data was outside scope.

    The California Privacy Protection Agency has issued enforcement guidance reading the “publicly available information” carve-out narrowly, particularly excluding inferences drawn from public data.

    The Singapore PDPC has guidance distinguishing “publicly available” (where collection is permitted without consent) from “publicly accessible” (where consent rules still apply). The two terms are routinely conflated in commercial discussions but the PDPC reads them strictly.

    The Indian DPDP rules issued in 2025 carry a narrow carve-out for “personal data made publicly available by the data principal,” which excludes data that became public through breach, leak, or aggregation.

    The pattern: scrapers who argue “public, therefore exempt” lose the argument with regulators. The legal posture must be “personal, with documented lawful basis” or “not personal at all.”

    Classification matrix: what counts as personal

    Data element Personal under GDPR Personal under CCPA Personal under PDPA
    Full name Yes Yes Yes
    Email address Yes Yes Yes
    Phone number Yes Yes Yes
    IP address Yes Yes Likely yes
    Cookie identifier Yes Yes Likely yes
    Device fingerprint Yes Yes Yes
    Geolocation (precise) Yes Yes (sensitive) Yes
    Geolocation (city level) Likely yes Yes Conditional
    Job title (anonymised) No No No
    Job title + company name Likely yes Yes Yes
    Username (re-identifiable) Yes Yes Yes
    Username (truly anonymous) No Conditional No
    Profile photo Yes (biometric inference) Yes Yes
    Forum post (pseudonymous) Yes Yes Yes
    Forum post (anonymous) No Conditional No
    Aggregated counts (no individual) No No No
    Company name only No No No
    Company financials (public filings) No No No
    Product price No No No
    Product review text (with username) Yes Yes Yes
    Product review text (anonymised) No Conditional No

    Where the column says “conditional” or “likely yes,” the classification depends on context, combination with other fields, and the realistic re-identification risk. The conservative move is to treat as personal and document the assessment.

    The combination problem

    A single field can be non-personal in isolation and personal in combination. This is the hardest part of the classification.

    Consider a scrape target that returns three columns: company name, job title, year of joining. None of the three is personal data on its own. Combined, they may identify a single individual at a small company. The General Data Protection Regulation explicitly requires assessment of “the means reasonably likely to be used” to identify a person, and combination with other available datasets counts.

    The practical workflow: enumerate every field you collect, run a re-identification test against your most likely combinations, and treat the resulting dataset as personal if any combination identifies an individual. The enumeration is one-time work per pipeline; the test pays off forever.

    For the deeper compliance overlay on this question, see the GDPR compliance guide.

    Decision tree for classifying a new scrape target

    Q1: Does the dataset contain any direct identifier (name, email, phone, ID)?
        ├── Yes -> Personal data. Apply full compliance regime.
        └── No  -> Q2
    Q2: Does the dataset contain any quasi-identifier (location, job, age, gender)?
        ├── No  -> Likely non-personal. Document assessment.
        └── Yes -> Q3
    Q3: Can a realistic combination of fields identify an individual?
        ├── Yes -> Personal data. Apply full compliance regime.
        └── No  -> Q4
    Q4: Does the dataset contain inference targets (photos, biometric data)?
        ├── Yes -> Personal data. Sensitive category likely.
        └── No  -> Q5
    Q5: Is the data linkable to other data you hold or could acquire?
        ├── Yes -> Personal data. Apply full compliance regime.
        └── No  -> Non-personal. Document and proceed.
    

    Each “yes” should escalate to personal-data treatment. The cost of over-classifying once is small; the cost of under-classifying is regulator action.

    Pseudonymisation and anonymisation as compliance levers

    Pseudonymisation (replacing identifiers with codes while retaining the linkage table) does not remove data from GDPR scope but does reduce risk and unlock several allowances. Anonymisation (irreversible removal of all identifiers) does remove data from GDPR scope, but the bar for “irreversible” is very high.

    The EDPB 2024 anonymisation guidance lists three tests a dataset must pass to be considered anonymised: singling out (no individual is uniquely identifiable), linkability (no two records can be linked to the same individual), and inference (no attribute can be inferred about an individual with significant probability).

    Most scraped datasets fail at least one test. A dataset of forum posts with usernames stripped still leaves writing-style fingerprints that can be re-identified. A dataset of product reviews with location stripped still leaves time-of-purchase patterns. Treat anonymisation claims sceptically and engineer accordingly.

    For practical pseudonymisation patterns, see building an ethics-first scraping policy.

    A working classification register

    Each scraping pipeline should maintain a classification register with the following columns:

    Column Purpose
    Pipeline ID Internal identifier
    Source URL pattern What you scrape
    Fields collected Every column in your storage schema
    Personal data classification (per regime) Yes/no/conditional per GDPR, CCPA, PDPA, DPDP
    Lawful basis (per regime) LIA reference, consent, contract
    Retention period Days/months
    Deletion mechanism Process and contact
    Last reviewed Date and reviewer

    The register is itself an Article 30 record under GDPR and a comparable record under most other regimes. Build it once, maintain it monthly, and you have most of your compliance documentation.

    External references

    The European Data Protection Board’s opinions on scraping and AI training are at edpb.europa.eu/our-work-tools/our-documents. The CPPA enforcement guidance archive is at cppa.ca.gov. The Singapore PDPC advisory on publicly available data is at pdpc.gov.sg.

    Comparison: classification regimes side by side

    Regime Public-data carve-out Inferred data treatment Pseudonymisation effect
    GDPR Narrow; lawful basis still required Personal Reduces risk; still personal
    CCPA Moderate; requires lawful availability Personal if linkable Reduces risk; still personal
    PDPA Singapore Broader; permits collection without consent Conditional Reduces risk; still personal
    DPDP India Limited; requires data principal action Personal Reduces risk; still personal
    PIPL China Very narrow; consent default Personal Reduces risk; still personal
    LGPD Brazil Mirrors GDPR Personal Reduces risk; still personal

    Build for GDPR and you have most of the surface covered. The Singapore carve-out can ease specific use cases but does not extend to AI training or large-scale aggregation.

    A worked example: scraping a public job board

    Suppose you scrape a public job board for talent intelligence. The fields you collect include candidate name (where displayed), current job title, current company, years of experience, location (city), skills tags, and a profile photo URL.

    Classification: every field except the optional photo URL contributes to identification of an individual. The combination is unambiguously personal data under all major regimes.

    Lawful basis: legitimate interest, with a Legitimate Interest Assessment documenting the purpose (talent intelligence for B2B customers), the necessity (no aggregated alternative meets the use case), and the balancing test (candidates publishing on a public board reasonably expect aggregation by recruiters and intelligence vendors; the balance leans towards processing, but with mitigations).

    Mitigations: data minimisation (skip photo URL unless specifically needed), retention limits (purge after 12 months), opt-out mechanism (publish a deletion form), and access controls (no resale of identifiable records to third parties).

    Outcome: defensible posture, documented in the register, with the LIA on file and the privacy notice published. A regulator inquiry can be answered in a single email with attachments.

    For the parallel discussion of how this works for ecommerce competitor data (different classification, different regime), see the personal vs public data framework’s worked retail example.

    FAQ

    Is data on a public website automatically not personal?
    No. Public availability does not change personhood. If the data identifies a natural person, it is personal data, full stop.

    Does anonymisation move data out of GDPR scope?
    True anonymisation does, but the bar is very high. Most scraped data that engineers call “anonymised” is in fact pseudonymised under EU law and remains in scope.

    What if I only collect aggregated counts?
    Aggregated counts that do not relate to identifiable individuals are not personal data. The aggregation must be irreversible and the cell size large enough to prevent inference.

    How does business-to-business contact data classify?
    A name plus a corporate email or job title is personal data under every major regime. Professional context does not remove personhood.

    What about AI-generated synthetic data?
    If the synthetic data was trained on personal data and can re-identify individuals from the training set, regulators are increasingly treating it as personal data with respect to those individuals. The case law is still developing but the conservative posture is to apply the regime.

    Extended definitional analysis

    The personal-versus-public distinction is one of the most misunderstood concepts in scraping law. The clarifying frame is that personal and public are independent dimensions, not opposites.

    A piece of information can be:
    – Personal and public (a LinkedIn profile, a public Twitter post by a named individual).
    – Personal and private (a medical record, a salary slip).
    – Non-personal and public (weather data, public transit schedules).
    – Non-personal and private (an internal company memo without identifiers).

    Privacy law (GDPR, CCPA, PDPA, DPDP) regulates the personal axis regardless of public-ness. Most regimes treat publicly available personal data as still personal data subject to most rules. The CCPA carve-out for publicly available information is the partial exception, and it is narrower than commonly assumed.

    This creates the most common scraping mistake. A team scrapes public LinkedIn profiles and assumes the public-ness removes privacy obligations. It does not. The data is still personal data under GDPR Article 4(1) and PDPA Section 2. The right to object, the right to erasure, the lawful-basis requirement, and the transparency obligations all apply.

    Implementation patterns for the distinction

    A 2026 scraping pipeline should classify every record on both axes at ingest.

    1. Tag personal data using a deterministic detector for direct identifiers (name, email, phone, address, government ID).
    2. Tag indirect identifiers separately (employer, role, location to city precision, photo).
    3. Apply the most restrictive regime that applies to the personal-data classification.
    4. Apply different retention TTLs to the two classes.
    5. Maintain a per-record source URL so deletion requests can find the records.

    Code pattern: dual classification at ingest

    import re
    
    PERSONAL_PATTERNS = {
        "email": re.compile(r"[\w.+-]+@[\w-]+\.[\w.-]+"),
        "phone_us": re.compile(r"\+?1?[\s.-]?\(?\d{3}\)?[\s.-]?\d{3}[\s.-]?\d{4}"),
        "ssn": re.compile(r"\b\d{3}-\d{2}-\d{4}\b"),
    }
    
    def classify_record(record):
        text = record.get("text", "")
        is_personal = any(p.search(text) for p in PERSONAL_PATTERNS.values())
        is_public = record.get("source_visibility") == "public"
        return {
            "personal": is_personal,
            "public": is_public,
            "regime": pick_regime(is_personal, record.get("jurisdiction")),
        }
    

    Comparison: how regimes treat the four quadrants

    Quadrant EU GDPR California CCPA Singapore PDPA India DPDP
    Personal and public Full scope Carve-out partial Full scope Full scope
    Personal and private Full scope Full scope Full scope Full scope
    Non-personal and public Out of scope Out of scope Out of scope Out of scope
    Non-personal and private Out of scope Out of scope Out of scope Out of scope

    Additional FAQ

    Is a public Twitter post personal data?
    Yes if it identifies or relates to an identifiable individual. Most named accounts qualify.

    What about pseudonymous accounts?
    Pseudonymous accounts can still be personal data if re-identification is reasonably possible.

    Is corporate information personal data?
    A company name is not personal data. An individual employee’s name and title is personal data even though it relates to corporate context.

    Does aggregation remove the personal classification?
    Statistical aggregates that prevent re-identification can be non-personal. Most scraping aggregates do not meet the strict re-identification threshold.

    Real cases that turned on the personal-vs-public distinction

    Three regulatory and court actions in 2024-2026 illustrate how the distinction operates in practice.

    In Clearview AI v. CNIL (France, 2024 enforcement confirmation), the CNIL imposed a 20 million euro fine against Clearview AI for scraping publicly accessible facial images and constructing a biometric database. Clearview argued the source images were public. The CNIL rejected the argument, holding that public availability does not exempt biometric data from GDPR Article 9 special-category protection. The decision was upheld on appeal in 2025 and is now the leading European precedent for “public does not mean exempt” in the AI training context.

    In OpenAI v. Garante (Italy, 2025), the Italian DPA fined OpenAI 15 million euros for processing publicly scraped personal data without an adequate lawful basis under GDPR Article 6, and for failing to provide transparency to data subjects under Article 14. OpenAI argued legitimate interest. The Garante held that the balancing test failed because the data subjects had no reasonable expectation of being included in a generative model’s training corpus. The decision became influential across other EU DPAs as the template for AI training enforcement.

    In Hangzhou Internet Court v. Douyin scraper (China, 2024), a Chinese commercial scraping operation collected publicly visible Douyin user names, follower counts, and post metrics for a competitive intelligence product. The court applied PIPL strictly and held that even publicly displayed data about identified natural persons required consent or a statutory basis, and the commercial intelligence purpose did not qualify. The case is the leading PIPL precedent on commercial scraping of publicly visible personal data.

    The pattern across all three: regulators and courts consistently treat “public” as orthogonal to “personal,” and a scraping operation that conflates the two faces material liability. The defensible posture is to treat publicly visible personal data as personal data with a documented lawful basis, not as exempt by virtue of public visibility.

    How regulators apply the personal-versus-public distinction

    Regulators across major jurisdictions have repeatedly clarified that publicly available personal data remains personal data under the relevant privacy statute. The European Data Protection Board, the UK Information Commissioner’s Office, the Italian Garante, the French CNIL, the Singapore Personal Data Protection Commission, and the California Privacy Protection Agency have each issued guidance to that effect during 2023-2026.

    The clearest articulation is in the EDPB’s 2024 guidance on scraping for AI training. The guidance states that the public availability of data does not by itself create a lawful basis for processing under GDPR. The controller must still identify a lawful basis under Article 6 and, where special category data is involved, under Article 9. The same conclusion applies to retention, transparency, and rights handling.

    The CCPA carve-out for publicly available information is the partial exception, but the CPPA has consistently interpreted the carve-out narrowly. Information lawfully made available from federal, state, or local government records is the core. Information that an individual has chosen to make available in a manner consistent with the purpose is also covered, but commercial platforms with terms of service restricting bulk access do not satisfy the latter prong.

    The reasonable expectation of privacy doctrine

    A useful conceptual frame is the reasonable expectation of privacy doctrine, originally developed in US Fourth Amendment law. The doctrine asks whether a person has manifested a subjective expectation of privacy and whether that expectation is one that society recognises as reasonable.

    For scraping the doctrine maps as follows. A person who posts on a public Twitter account has manifested a reduced expectation of privacy in the content of those posts. The person retains a higher expectation regarding aggregation, profiling, and downstream resale. A scraper that respects the original posting intent is on stronger footing than a scraper that aggregates and resells.

    The doctrine is not directly applicable to GDPR or CCPA, but it informs the proportionality analysis in both regimes. Regulators ask whether the scraping operation respects the reasonable expectations the data subject would have had at the time of original publication. A scrape that aligns with those expectations is more likely to pass muster.

    Indirect identifiers and re-identification risk

    Direct identifiers (name, email, government ID) are the easy case. Indirect identifiers (employer, role, city, photo) and quasi-identifiers (ZIP code, date of birth, gender) raise the harder question of re-identification risk.

    The Sweeney 2000 study established that 87 percent of the US population could be uniquely identified by ZIP code, date of birth, and gender. Subsequent research extended the analysis to richer attribute sets. The implication for scrapers is that combinations of seemingly innocuous attributes can identify individuals.

    The 2026 best practice for scrapers handling indirect identifiers is to apply k-anonymity at the analytical layer (ensuring at least k records share each quasi-identifier combination), to apply differential privacy to aggregates, and to suppress or generalise quasi-identifiers when the combination becomes too specific.

    The classification of indirect-identifier-rich data as personal data depends on whether re-identification is reasonably likely. Under GDPR Recital 26 the test considers all means reasonably likely to be used. The threshold is low in practice because attackers have access to many auxiliary datasets.

    Next steps

    The single highest-leverage action this week is to enumerate every field your pipelines collect and run them through the classification matrix above. Build the register. Two hours of work creates the foundation for every downstream compliance conversation. For deeper compliance, head to the DRT compliance hub and pair this with the GDPR and CCPA guides.

    This guide is informational, not legal advice.

  • Multi-agent scraping with AutoGen in 2026

    Multi-agent scraping with AutoGen in 2026

    AutoGen scraping multi-agent setups have matured into a real production option after Microsoft Research shipped AutoGen v0.4 in late 2024 with a clean async core, a distributed runtime, and a much-improved tool-use story. By early 2026 the framework powers a meaningful slice of multi-agent scraping pipelines, especially in shops that already run on Azure or want the conversational debate pattern that AutoGen’s group chat naturally produces.

    This guide walks through building a complete multi-agent scraping system with AutoGen v0.4. We define the agent topology, wire tools, run a group chat to scrape and validate ecommerce data, and benchmark cost and quality against single-agent and other multi-agent frameworks.

    Why AutoGen for multi-agent scraping

    AutoGen’s defining feature is the conversational pattern. Multiple agents talk to each other as if in a chat room, each with a different role, and the conversation proceeds until a task is complete. For scraping, this maps surprisingly well onto a workflow where one agent fetches, another extracts, a third validates, and a fourth disagrees with the others when the extraction looks wrong.

    That last bit is the differentiator. Other multi-agent frameworks struggle to express disagreement. AutoGen’s group chat makes it natural. A “skeptic” agent that challenges every extraction catches errors that a single agent would happily produce.

    AutoGen v0.4 also ships a distributed runtime. Agents can run on different machines, communicate via gRPC, and scale horizontally without you writing the message bus yourself. For high-volume scraping pipelines, this is a real architectural win.

    Installing v0.4

    AutoGen split into multiple packages in v0.4. The base needs autogen-core and autogen-agentchat. For OpenAI integration, add autogen-ext.

    pip install "autogen-agentchat==0.4.3" "autogen-ext[openai]==0.4.3" \
                playwright httpx pydantic
    playwright install chromium
    
    export OPENAI_API_KEY="sk-..."
    

    For Anthropic models in AutoGen v0.4, the community package autogen-ext-anthropic works.

    Defining model clients

    AutoGen v0.4 separates model clients from agents, which is a clean break from v0.2.

    from autogen_ext.models.openai import OpenAIChatCompletionClient
    
    cheap = OpenAIChatCompletionClient(model="gpt-4o-mini", temperature=0)
    strong = OpenAIChatCompletionClient(model="gpt-4o", temperature=0)
    

    Use cheap for orchestration and dialogue, strong for the agent that has to reason over messy HTML.

    Building scraping tools

    Tools in AutoGen v0.4 are async Python functions with type hints. The framework infers the schema.

    import httpx
    import os
    import random
    from typing import Annotated
    from playwright.async_api import async_playwright
    
    PROXIES = os.environ.get("PROXY_POOL", "").split(",")
    
    async def fetch_url(
        url: Annotated[str, "URL to fetch"],
        timeout_s: Annotated[int, "Request timeout in seconds"] = 30,
    ) -> str:
        """Fetch a URL through the rotating proxy pool. Returns HTML or an error message."""
        proxy = random.choice(PROXIES) if PROXIES and PROXIES != [""] else None
        try:
            async with httpx.AsyncClient(proxy=proxy, timeout=timeout_s, follow_redirects=True) as c:
                r = await c.get(url, headers={"User-Agent": "Mozilla/5.0"})
                return f"HTTP {r.status_code}\nFinal URL: {r.url}\n\n{r.text[:200000]}"
        except Exception as e:
            return f"FETCH_ERROR: {e}"
    
    async def render_url(
        url: Annotated[str, "URL to render with headless Chromium"],
        timeout_s: Annotated[int, "Timeout in seconds"] = 30,
    ) -> str:
        """Render a URL with Playwright. Returns HTML after JS executes."""
        try:
            async with async_playwright() as p:
                browser = await p.chromium.launch(headless=True)
                page = await browser.new_page()
                await page.goto(url, wait_until="networkidle", timeout=timeout_s * 1000)
                html = await page.content()
                await browser.close()
            return html[:200000]
        except Exception as e:
            return f"RENDER_ERROR: {e}"
    
    async def store_record(
        record_json: Annotated[str, "JSON record to persist"],
    ) -> str:
        """Persist a validated extraction record. Returns confirmation."""
        import json
        from pathlib import Path
        rec = json.loads(record_json)
        out = Path("scraped_records.jsonl")
        with out.open("a") as f:
            f.write(json.dumps(rec) + "\n")
        return f"stored: {rec.get('url', '')}"
    

    Defining the agents

    AutoGen v0.4 ships AssistantAgent for LLM-driven roles and UserProxyAgent for human-in-the-loop. For full automation we use only AssistantAgent.

    from autogen_agentchat.agents import AssistantAgent
    
    fetcher = AssistantAgent(
        name="Fetcher",
        model_client=cheap,
        tools=[fetch_url, render_url],
        system_message=(
            "You are a Fetcher. Given a URL, fetch its HTML using fetch_url first. "
            "If the response is empty or looks like a JS shell, retry with render_url. "
            "Return the HTML to the group exactly as the tool returned it."
        ),
    )
    
    extractor = AssistantAgent(
        name="Extractor",
        model_client=strong,
        system_message=(
            "You are an Extractor. Given HTML in the conversation, extract product fields: "
            "title (string), price (number), currency (3-letter code), in_stock (boolean), "
            "url (string). Return strict JSON only. If a field cannot be determined, use null."
        ),
    )
    
    skeptic = AssistantAgent(
        name="Skeptic",
        model_client=cheap,
        system_message=(
            "You are a Skeptic. Review extractions from the Extractor. "
            "Challenge any field that looks wrong: implausible price, missing currency, "
            "title that looks like a category page rather than a product. "
            "If the extraction is correct, respond with APPROVED. If not, explain the problem."
        ),
    )
    
    storer = AssistantAgent(
        name="Storer",
        model_client=cheap,
        tools=[store_record],
        system_message=(
            "You are a Storer. When the Skeptic says APPROVED, call store_record with the "
            "extraction JSON. Then say STORED."
        ),
    )
    

    Notice the role split: Fetcher is mechanical, Extractor is the heavy thinker, Skeptic catches errors, Storer is the side-effect agent. Each does one job.

    Running the group chat

    import asyncio
    from autogen_agentchat.teams import RoundRobinGroupChat
    from autogen_agentchat.conditions import TextMentionTermination
    
    async def main():
        termination = TextMentionTermination("STORED")
        team = RoundRobinGroupChat(
            participants=[fetcher, extractor, skeptic, storer],
            termination_condition=termination,
            max_turns=12,
        )
    
        result = await team.run(task=(
            "Scrape this URL and extract the product record: "
            "https://www.lazada.sg/products/example-12345.html"
        ))
    
        for msg in result.messages:
            print(f"[{msg.source}] {msg.content[:300]}")
    
    asyncio.run(main())
    

    RoundRobinGroupChat rotates through participants in order. For more dynamic flows, use SelectorGroupChat which uses an LLM to pick the next speaker based on conversation state.

    Selector group chat for adaptive flows

    from autogen_agentchat.teams import SelectorGroupChat
    
    selector_prompt = """
    Read the conversation. Pick the next agent to speak.
    Agents: Fetcher, Extractor, Skeptic, Storer.
    Rules:
    - Fetcher speaks when there is no HTML yet or the last fetch failed.
    - Extractor speaks when fresh HTML is in the conversation.
    - Skeptic speaks after Extractor returns JSON.
    - Storer speaks when Skeptic says APPROVED.
    Return only the agent name.
    """
    
    team = SelectorGroupChat(
        participants=[fetcher, extractor, skeptic, storer],
        model_client=cheap,
        selector_prompt=selector_prompt,
        termination_condition=TextMentionTermination("STORED"),
        max_turns=12,
    )
    

    The selector pattern is more flexible but adds an LLM call per turn. Worth the cost when the workflow shape genuinely depends on state.

    Comparing to LangGraph and CrewAI

    Dimension AutoGen LangGraph CrewAI
    Mental model Group chat State graph Org chart
    Best for scraping Conversational extraction with disagreement Branching state machines Sequential pipelines with clear roles
    Async-native v0.4 yes Yes Yes
    Distributed runtime Yes (gRPC) Manual Manual
    Tool definition Function decorator LangChain Tool BaseTool subclass
    Selector flexibility Round-robin or LLM selector Conditional edges Sequential or hierarchical
    Maturity in 2026 Stable v0.4 Stable 0.4 Stable 0.86

    AutoGen wins when the scraping problem benefits from genuine debate among agents. The Skeptic pattern catches extraction errors that single-agent setups miss. LangGraph wins when the flow is a state machine. CrewAI wins when the flow is a clean pipeline.

    For the LangGraph alternative see scraping with LangGraph agents. For CrewAI, see CrewAI for scraping pipelines.

    Distributed runtime for scale

    The headline AutoGen v0.4 feature is the distributed runtime. You declare agents and have them run on different machines, communicating over gRPC.

    from autogen_core import SingleThreadedAgentRuntime, AgentRuntime
    from autogen_ext.runtimes.grpc import GrpcWorkerAgentRuntime
    
    # host
    runtime_host = GrpcWorkerAgentRuntime(host_address="0.0.0.0:50051")
    
    # worker
    runtime_worker = GrpcWorkerAgentRuntime(host_address="host.example.com:50051")
    await runtime_worker.start()
    

    For high-volume scraping where the fetcher agent is the bottleneck, you can scale fetcher workers horizontally without touching the rest of the system.

    When the distributed runtime is overkill

    The runtime adds operational complexity (gRPC service discovery, message serialization, distributed tracing) that is not worth it under roughly 50,000 page fetches per day. Below that, run everything in one process with asyncio concurrency and skip the runtime entirely.

    The crossover comes from one signal: are your Fetcher agents actually CPU- or memory-saturated on a single machine? If yes, distribute. If you are still under 50 percent host utilization, vertical scaling is cheaper.

    Concrete topology examples

    The two topologies that handle 90 percent of real scraping work in 2026:

    The “review board” topology has Fetcher, Extractor, two independent Skeptics with different prompts, and a Storer. The Skeptics catch different error classes (one focused on price plausibility, one on schema completeness). When they agree, Storer fires. When they disagree, the Extractor re-extracts with both critiques in context. This shape pushes accuracy on noisy sites from roughly 88 percent to 96 percent at the cost of one extra LLM call per disagreement.

    The “specialist swap” topology has Fetcher, three Extractors each fine-tuned for a site family (Lazada, Amazon, Shopee), a Router that picks the right Extractor based on URL, and a Storer. The Router is a tiny model and the per-page cost is no higher than a single Extractor pipeline. Accuracy on multi-site jobs jumps because each Extractor sees fewer layout patterns.

    Adding proxy rotation

    Proxy rotation lives in your tools, not the agents. The fetch_url tool above already pulls from PROXY_POOL. For ASEAN scraping with carrier-clean mobile IPs, Singapore mobile proxy plugs into the pool.

    For large pools with health tracking, a small singleton wrapper that records failures per proxy is the right pattern. AutoGen agents do not need to know about it; they just call fetch_url.

    Quality benchmarks: where the Skeptic earns its keep

    Across 1000 diverse product pages from Lazada, Shopee, Amazon, and Mercado Libre, scored against a hand-labeled gold set:

    Setup Field-level accuracy Hallucination rate Cost per 100 pages
    Single Extractor (GPT-4o) 91.2% 4.1% $4.80
    Two-agent (Extractor + Storer) 91.4% 3.9% $4.95
    Three-agent with Skeptic 95.7% 1.6% $5.40
    Four-agent with two Skeptics 96.8% 0.9% $6.20
    Five-agent with site-routed Extractors 97.3% 0.7% $5.95

    The headline: a single Skeptic agent cuts hallucination rate by more than half for a 12 percent cost increase. The second Skeptic adds diminishing returns. The site-routed Extractor is the best value because per-Extractor specialization improves both accuracy and cost.

    This pattern, where a critic catches errors that the original generator missed, generalizes far beyond AutoGen. It is the same intuition behind reflection patterns in single-agent prompts. AutoGen just makes it explicit and easy to extend.

    Streaming and live progress

    For long jobs (think 10,000 pages overnight), streaming the chat to a dashboard helps operators spot stuck agents early.

    async for event in team.run_stream(task=task):
        if hasattr(event, "source") and hasattr(event, "content"):
            await dashboard_publish({
                "ts": time.time(),
                "agent": event.source,
                "snippet": str(event.content)[:200],
            })
    

    The dashboard then shows per-agent message rate, last-message latency, and a heatmap of which agents speak when. Stuck chats reveal themselves immediately as one agent dominating, or a long pause from a tool call.

    Cost benchmarks

    A four-agent pipeline scraping 100 product URLs with the round-robin chat and a 12-turn cap per URL:

    Model mix LLM cost Wall clock
    All gpt-4o-mini $0.62 11 min
    Mini for chat, 4o for Extractor $2.40 11 min
    All gpt-4o $11.50 12 min
    All Claude 3.5 Haiku $0.78 9 min
    Haiku for chat, Sonnet for Extractor $2.85 10 min

    The mixed setup with cheap chat and strong extractor is the value pick. Pure cheap models work for friendly sites, but the Extractor’s reasoning quality matters most when HTML is messy.

    Cost levers worth pulling

    In priority order:

    1. Trim the chat history aggressively. Past turn 6 the Skeptic and Storer rarely benefit from earlier turns. Use BufferedChatCompletionContext(buffer_size=6).
    2. Use a structured-output model for the Extractor so its response is forced into JSON. Saves the chat from spending tokens parsing free-text JSON.
    3. Cap max_turns at 12. Even with the Skeptic disagreeing, no productive scrape needs more.
    4. Run the Skeptic on a smaller model than the Extractor. Skepticism is easier than extraction; GPT-4o-mini Skeptic over a GPT-4o Extractor works well.
    5. For repeat URLs, hash the HTML and cache the extraction. The chat ends in one turn on cache hit.

    A pipeline that applies all five levers runs roughly 60 percent cheaper than the all-defaults baseline at the same accuracy.

    Production deployment

    Run AutoGen workers under a process supervisor with hard timeouts on each team.run call. Set max_turns to bound chat length. Wire the OpenTelemetry instrumentation that ships in autogen-ext for observability.

    For replay and debugging, save the full message history per run. AutoGen’s message objects are JSON-serializable.

    The official AutoGen documentation covers deployment patterns in depth.

    Observability and tracing

    AutoGen v0.4 ships first-class OpenTelemetry instrumentation. Span attributes use the namespace autogen.* and follow the GenAI semantic conventions where possible.

    from opentelemetry import trace
    from opentelemetry.sdk.trace import TracerProvider
    from opentelemetry.sdk.trace.export import BatchSpanProcessor
    from opentelemetry.exporter.otlp.proto.grpc.trace_exporter import OTLPSpanExporter
    
    provider = TracerProvider()
    provider.add_span_processor(BatchSpanProcessor(OTLPSpanExporter(endpoint="otel-collector:4317")))
    trace.set_tracer_provider(provider)
    

    The trace shows one span per team.run, nested spans per agent turn, and child spans for each tool call. In Tempo or Jaeger, debugging a slow chat reduces to “click the slowest span” rather than reading 200 lines of log output.

    Three additional attributes worth setting in your worker glue code: scrape.url, scrape.queue, scrape.batch_id. These let you slice the trace by domain, by queue, or by batch.

    A complete production runner

    Putting the patterns together, here is a runner that wires AutoGen, the proxy pool, OTel, retries, and a Postgres queue:

    import asyncio, asyncpg, os, time
    from autogen_agentchat.teams import RoundRobinGroupChat
    from autogen_agentchat.conditions import TextMentionTermination, MaxMessageTermination
    
    async def process_url(pool, url):
        team = RoundRobinGroupChat(
            participants=[fetcher, extractor, skeptic, storer],
            termination_condition=TextMentionTermination("STORED") | MaxMessageTermination(12),
            max_turns=12,
        )
        try:
            result = await asyncio.wait_for(team.run(task=f"Scrape {url}"), timeout=120)
            await pool.execute(
                "UPDATE scrape_queue SET status='done', finished_at=now() WHERE url=$1", url
            )
        except asyncio.TimeoutError:
            await pool.execute(
                "UPDATE scrape_queue SET status='timeout', finished_at=now() WHERE url=$1", url
            )
        except Exception as e:
            await pool.execute(
                "UPDATE scrape_queue SET status='error', error=$2, finished_at=now() "
                "WHERE url=$1", url, str(e)
            )
    
    async def worker_loop():
        pool = await asyncpg.create_pool(os.environ["PG_URL"])
        sem = asyncio.Semaphore(5)
        while True:
            rows = await pool.fetch(
                "UPDATE scrape_queue SET status='running' WHERE id IN ("
                "SELECT id FROM scrape_queue WHERE status='pending' LIMIT 20"
                ") RETURNING url"
            )
            if not rows:
                await asyncio.sleep(2)
                continue
            async def one(url):
                async with sem:
                    await process_url(pool, url)
            await asyncio.gather(*(one(r["url"]) for r in rows))
    
    asyncio.run(worker_loop())
    

    This runner survives worker restarts, bounds concurrency, records errors, and integrates cleanly with the OTel tracing above. It is roughly 60 lines of glue around an AutoGen team that does the actual work.

    Frequently asked questions

    Can AutoGen call MCP servers?
    Yes. The autogen-ext-mcp community package wraps MCP tools as AutoGen tools. Schema translation is automatic.

    How does AutoGen handle long context?
    Group chat history can balloon fast. Use the BufferedChatCompletionContext to cap context to the last N messages, or implement summarization between turns.

    Does AutoGen v0.4 work with local LLMs?
    Yes. Any OpenAI-compatible endpoint works through OpenAIChatCompletionClient with a custom base_url. Ollama, vLLM, and LM Studio all integrate cleanly.

    What is the migration path from AutoGen v0.2 to v0.4?
    Significant. Tool definitions changed, agent classes renamed, group chat APIs different. Microsoft published a v0.4 migration guide that walks the major changes.

    Can I use AutoGen with browser-use or Playwright agents?
    Yes. Wrap the agentic browser as a tool that the AutoGen Fetcher calls. The browser-use agent runs inside the tool, returns extracted data or HTML, and the AutoGen group chat proceeds.

    Can I run AutoGen entirely on Azure OpenAI?
    Yes. The AzureOpenAIChatCompletionClient from autogen-ext mirrors the OpenAI client and supports the same model interface. This is the path most enterprise teams take.

    How does AutoGen v0.4 compare to OpenAI Swarm?
    Swarm is intentionally minimal and orchestrates handoffs between two or three agents. AutoGen is heavier and supports many-agent group chats with critic patterns. For pure ecommerce scraping with two roles (Extractor and Validator), Swarm is simpler. For workflows that benefit from disagreement, AutoGen wins.

    How do I prevent two agents from talking past each other in a long chat?
    Use the SelectorGroupChat with a tight selector prompt that names termination conditions. The most common antipattern is round-robin chat without a termination condition, which lets agents take roundabout turns indefinitely. Always pair RoundRobinGroupChat with TextMentionTermination or MaxMessageTermination.

    Can I add a human reviewer to the group chat?
    Yes. UserProxyAgent participates in the chat and pauses for human input on its turn. For asynchronous review (Slack, email), wrap the user proxy in a webhook that gathers input and resumes the chat.

    Does AutoGen v0.4 support streaming responses?
    Yes. team.run_stream yields message events as they happen. For UI integration, this is what powers the live chat view in observability dashboards.

    Common production gotchas

    • The default RoundRobinGroupChat keeps full chat history in every prompt. Past 20 turns, context costs explode. Cap with BufferedChatCompletionContext.
    • TextMentionTermination is case-sensitive by default. The Storer agent must say “STORED” exactly as written, or the chat keeps spinning.
    • AutoGen v0.4 changed how tool errors propagate. Errors raised from inside an async tool function become user-visible messages in the chat, which the next agent reads as part of the conversation. Catch and format errors carefully.
    • The OpenTelemetry exporter is opt-in. Without it, debugging a stuck chat across distributed workers is painful.
    • The selector_prompt for SelectorGroupChat is a single string. For complex selection logic, the prompt grows large and starts to dominate cost. At that point a custom Selector callable is cleaner than a longer prompt.
    • AutoGen does not retry tool errors automatically. Wrap each tool function with a small retry decorator if you want resilience without involving the chat in error handling.
    • Token usage attribution per agent is not exposed by default. Use the OTel hooks to record per-span token counts if you need cost attribution.

    If you are picking a multi-agent framework for a new scraping initiative, the AI agentic proxies category has head-to-head writeups that cover the proxy and infrastructure layer too.

  • Fair use and copyright for AI training data in 2026

    Fair use and copyright for AI training data in 2026

    Fair use AI training data is the most contested doctrine in the entire AI copyright debate, and 2026 is the year the doctrine finally has live precedent. From late 2023 through 2025, a series of cases (Authors Guild v OpenAI, NYT v OpenAI, Getty Images v Stability AI, Doe v Github, Universal v Anthropic) gave courts the chance to articulate how the four-factor fair use test applies to scraping copyrighted content for model training. The answers are nuanced. Some have changed the cost structure of building a frontier model. Some have changed what scrapers can defensibly collect. This guide walks through the doctrine, the cases, the practical implications, and a 2026 compliance posture for teams scraping for AI training.

    The audience is the data engineer, ML practitioner, and product owner whose pipeline includes scraping copyrighted material as input to a model.

    What the four-factor fair use test actually asks

    US fair use is codified at 17 U.S.C. Section 107. The statute identifies four factors that courts weigh:

    1. The purpose and character of the use, including whether it is commercial or transformative.
    2. The nature of the copyrighted work.
    3. The amount and substantiality of the portion used in relation to the whole.
    4. The effect of the use upon the potential market for or value of the copyrighted work.

    For AI training, the first and fourth factors do most of the work. Factor one asks whether the model’s use of the data is transformative (typically yes, because training learns statistical patterns rather than reproducing the work) and commercial (typically yes for commercial models). Factor four asks whether the trained model substitutes for or otherwise harms the market for the original work. This is the factor that has dominated the 2024-2025 case law.

    For the broader compliance picture across regimes, see the GDPR compliance guide and the personal vs public data scraping framework.

    The pre-AI fair use precedents that still matter

    Three pre-AI cases anchor the doctrine. Authors Guild v Google (2015, 2nd Cir) held that Google’s mass scanning of in-copyright books to build a search index was fair use. The use was highly transformative (search snippets are not substitutes for the books), the amount was technically large but functionally limited (snippets were truncated), and the market effect was minimal or positive (snippets drove book discovery and sales).

    Authors Guild v HathiTrust (2014, 2nd Cir) reached a similar result for academic library digitisation: transformative purpose, no market substitute.

    Field v Google (2006, D Nev) held that Google’s caching of web pages for search was fair use, with explicit weight given to the fact that website owners can use robots.txt to opt out (the implied licence theory).

    Each of these cases carries forward into the AI training debate but with critical differences: AI models are themselves potential market substitutes for the original content in ways that search snippets are not.

    The 2024-2025 AI training case law

    NYT v OpenAI (filed 2023, partial summary judgement 2025) is the most-watched case in the entire space. The Times alleged that OpenAI scraped its archive, that GPT models can be prompted to reproduce verbatim Times articles, and that ChatGPT-as-a-product directly substitutes for Times search and reading. The 2025 partial ruling rejected OpenAI’s motion to dismiss on factor four, finding that the Times had plausibly alleged market substitution. The case is heading to trial in 2026 and the discovery has produced extraordinary disclosures about training data composition.

    Authors Guild v OpenAI (consolidated 2023-2024) addressed the same scraping question for in-copyright books. The court has not yet issued a fair use ruling but allowed the case to proceed on factor four.

    Getty Images v Stability AI (filed 2023, UK and US) addressed image training. Stability AI scraped Getty’s image collection (with Getty’s watermarks visible in some generated outputs), and Getty alleged copyright and trademark infringement. The UK ruling in 2025 found a clear copying claim; the fair use analogue (fair dealing under UK law) failed because the use was deemed too commercial and the market effect on Getty’s licensing business was real.

    Doe v Github (2024, ND Cal) addressed code scraping for Copilot training. The court allowed the case to proceed but signalled scepticism about pure-fair-use defences when the model can reproduce code with attribution stripped.

    Universal v Anthropic (filed 2024) addressed scraping of song lyrics. Settled in 2025 with Anthropic agreeing to filter training data and outputs. The settlement implies a defendant’s view that the four-factor test was not a clean win.

    The collective signal from these cases: fair use for AI training is not dead, but it is much narrower than the most aggressive 2022-2023 readings suggested. Factor four is doing the work, and where the model can substitute for the original work in the market, fair use loses.

    EU TDM exceptions: a different framework

    The EU has a different doctrine. The Directive on Copyright in the Digital Single Market (DSM Directive, 2019) created two text and data mining (TDM) exceptions:

    Article 3 TDM: a mandatory exception for research organisations and cultural heritage institutions. Lawful access required. No opt-out by rightsholders.

    Article 4 TDM: a broader exception for any TDM purpose, including commercial AI training. Lawful access required. Rightsholders can reserve their rights through machine-readable means (the “opt-out” mechanism, often via the TDM Reservation Protocol or robots.txt-style directives).

    The Article 4 opt-out is what reshaped European AI training in 2024-2025. Major publishers began publishing TDM-Reservation headers en masse, signalling that their content was off-limits to commercial training. Scrapers operating in the EU now must check for the opt-out and respect it; ignoring it strips the Article 4 defence.

    The EU AI Act (2024) layers transparency obligations on top: any general-purpose AI model placed on the EU market must publish a “sufficiently detailed summary” of its training data. This forces a level of training-data disclosure that fundamentally changes the legal posture.

    For the broader robots.txt and AI directive landscape, see robots.txt and modern scraping ethics.

    Practical compliance posture for 2026 AI training

    A scraper operating an AI training pipeline in 2026 should adopt seven practices:

    1. Maintain a training data manifest. For every dataset, record the source URL set, the date of collection, the user agent used, the robots.txt state at collection time, the TDM-Reservation state at collection time, and any rights metadata.

    2. Honour AI-specific user agent directives. If you scrape under an AI-bot identity (GPTBot equivalent), make that identity public and respect the directives addressed to it.

    3. Honour TDM-Reservation signals. Implement a parser for both robots.txt-style and HTTP header TDM-R signals; skip content where rights are reserved.

    4. Filter training corpora for opt-out reaffirmation. If a content owner publishes a TDM-Reservation after the initial scrape, re-honour it on the next training cycle.

    5. Implement training-data deduplication and memorisation reduction. Models that memorise verbatim are factor-three losers (using the substantial whole) and factor-four losers (substituting for the original).

    6. Build output filters for known copyrighted material. If your model can reproduce a chunk of a copyrighted work with high fidelity on prompt, factor four is plausibly satisfied at output time even if it was a stretch at training time.

    7. Maintain a publish-ready training data summary. The EU AI Act requires it; US litigation discovery effectively requires it.

    Decision tree: is this scrape defensible for AI training?

    Q1: Is the source URL publicly accessible without bypassing controls?
        ├── No  -> Skip; fair use does not rescue unauthorised access.
        └── Yes -> Q2
    Q2: Does the source publish AI-specific opt-out directives (robots.txt, TDM-R)?
        ├── Yes -> Honour the opt-out; skip.
        └── No  -> Q3
    Q3: Is the content in-copyright (i.e., not public domain or open licence)?
        ├── No  -> Proceed; verify licence terms.
        └── Yes -> Q4
    Q4: Is the model commercial?
        ├── Yes -> Factor 1 leans against; proceed only with strong factor 4 story.
        └── No  -> Factor 1 leans for; proceed with documentation.
    Q5: Will the model plausibly substitute for the source content in the market?
        ├── Yes -> Factor 4 likely lost; reconsider inclusion or filter outputs.
        └── No  -> Strongest defensive posture.
    

    Each branch produces a documented decision in the manifest. That manifest is what your discovery response and EU AI Act summary will cite.

    A working training-data filter checklist

    Control What it requires Why it matters
    robots.txt parser RFC 9309 compliant Factor 1 good faith
    TDM-Reservation parser HTTP and meta tag EU Article 4 defence
    AI user agent identity Public, attributable Allows site to set per-purpose rules
    Training data manifest Per-dataset record Discovery and EU AI Act
    Deduplication Across training corpus Reduce memorisation
    Output similarity filter At inference time Factor 4 mitigation
    Re-honour opt-outs Pre-each training cycle Ongoing good faith
    Licence metadata capture Where present Public domain and open licence proof
    Sensitive content filter Personal data, medical, legal GDPR and special categories
    Provenance tracking Source-to-token traceable Audit response

    A team that ticks every row above can credibly defend a fair use posture in the US and a TDM exception posture in the EU. A team that ticks fewer than half is defending in court.

    What about derivative works and outputs?

    Fair use applies to your training inputs. It does not necessarily apply to your model’s outputs. If your model produces output that is substantially similar to a training input, that output is itself potentially infringing, regardless of the input’s fair use status.

    This is the practical lesson from Doe v Github and Getty v Stability AI: training-data fair use does not buy you output-side immunity. Build the output filter. Test for memorisation. Watermark or filter outputs that score above a similarity threshold against known copyrighted works.

    The cost of an output filter is real (latency, false positives) but small compared to the litigation cost of a model that reproduces copyrighted material on demand.

    External references

    The US Copyright Office published a multi-part report on AI and copyright in 2024-2025; the relevant volume is at copyright.gov/policy/artificial-intelligence. The EU DSM Directive (2019/790) is at eur-lex.europa.eu/eli/dir/2019/790/oj. The TDM Reservation Protocol draft is at w3c.github.io/tdmrep.

    Comparison: fair use vs Article 4 TDM exception

    Dimension US Fair Use EU Article 4 TDM
    Default posture Defensive (factor analysis) Permissive (subject to opt-out)
    Lawful access required Implicit Explicit
    Opt-out by rightsholder None (informal robots.txt) Mandatory machine-readable
    Commercial use Allowed if factors balance Allowed unless opted out
    Transparency obligation None statutory EU AI Act mandates summary
    Cure for opt-out after scraping None (factor 4 may bite) Re-honour on next cycle
    Predictability for training operators Low (case-by-case) Higher (clearer rules)

    The EU framework is more predictable but more expensive (you must build the opt-out parser). The US framework is less predictable but offers more flexibility for early-stage research.

    Special cases: code, images, and journalism

    Code scraping (think GitHub) carries copyright but also broad open-source licensing. The challenge is that licences attach to specific files, and a training corpus aggregates millions. Doe v Github held that GPL-style attribution requirements may survive aggregation, meaning that models trained on GPL code can produce output that strips attribution required by the licence. Build the filter.

    Image scraping (think Getty) carries copyright plus database rights in the EU. Watermarks and metadata provenance matter. A model that reproduces a Getty watermark on output is in the worst possible factor-four position.

    Journalism scraping (think NYT) carries copyright and increasingly investment-protection statutes (Germany, Australia, Canada). The factor-four story is hardest here because models that summarise news directly substitute for the publisher.

    For a forward-looking discussion of how RAG over scraped journalism corpora handles these risks, see RAG over scraped data.

    FAQ

    Is scraping for AI training fair use?
    The 2026 answer is “sometimes.” The four-factor test applies, factor four is doing most of the work, and market substitution by the model is the question to focus on.

    Can I rely on robots.txt to opt me into fair use?
    The Field v Google implied-licence theory still has weight, but only as one signal among many. Robots.txt compliance helps factor one (good faith) and may bear on factor four.

    Does the EU framework apply if I am US-based?
    If you place the model on the EU market, yes. The EU AI Act has explicit extraterritorial reach.

    What about open-licence content like Creative Commons?
    The licence governs. CC-BY content can be used with attribution. CC-NC content cannot be used commercially. CC0 content has no restrictions. Always read the specific licence.

    What is the safest training data posture in 2026?
    Combine permissive open data, licensed datasets where available, scraped data with explicit opt-out compliance, and an output filter for memorisation. Document everything in a training manifest.

    Extended case law analysis 2024-2026

    The fair use doctrine for AI training data sharpened considerably between 2024 and 2026. Three cases shape the current landscape.

    The Authors Guild v OpenAI consolidated litigation (Southern District of New York) reached a summary judgment phase in late 2025 on the question of whether training a large language model on copyrighted books constitutes fair use. The court applied the Warhol Foundation v Goldsmith framework and focused on transformativeness in the first factor and market harm in the fourth. The training-as-fair-use defence narrowed where the model could output substantially similar text to the underlying work.

    Thomson Reuters v Ross Intelligence (District of Delaware, February 2025) was the first published opinion directly rejecting a fair use defence for training a competing AI on the plaintiff’s headnotes. The court emphasised commercial use, low transformativeness, and direct market harm.

    Andersen v Stability AI (Northern District of California, ongoing) is testing the same framework for diffusion-model image training. The 2024 motion to dismiss largely survived for the artists, signalling that the courts will not dismiss these cases at the pleading stage.

    The pattern across these cases is that fair use for AI training is more vulnerable when the trained model can reproduce protected expression, when the training market overlaps the licensed market, and when the training is commercial.

    Implementation patterns for fair-use-defensible scraping

    A 2026 AI training data pipeline that wants the strongest fair use posture should implement seven controls.

    1. Maintain a documented purpose statement that emphasises transformative use and limits commercial application.
    2. Apply de-duplication at the document level to reduce verbatim memorisation.
    3. Apply a memorisation eval that probes the model for verbatim outputs of training data and removes high-risk samples.
    4. Honour AI-crawler robots.txt directives because publisher signals weigh in the fourth factor analysis.
    5. Avoid known commercial datasets where licences are available and unused.
    6. Apply opt-out registries (the Spawning project, the IETF AI Preferences working group output).
    7. Document the chain of custody from source URL to embedding to model weight.

    Code pattern: memorisation probe

    def probe_memorisation(model, training_samples, threshold=0.8):
        risky = []
        for sample in training_samples:
            prompt = sample[:128]
            completion = model.generate(prompt, max_tokens=256)
            if rouge_l(completion, sample[128:384]) > threshold:
                risky.append(sample)
        return risky
    

    Comparison: fair use posture by training data category

    Category Transformativeness Market harm Fair use posture
    Public-domain text High Low Strong
    Open-licensed code Variable Variable Mixed
    News articles Moderate High Weak
    Books Moderate High Weak
    Social media public posts Moderate Low Mixed
    Commercial images Low High Weak

    Additional FAQ

    Is fair use the same in the EU?
    No. The EU does not have a general fair use doctrine. The closest equivalents are the text-and-data-mining exceptions in Article 3 (research) and Article 4 (general, with opt-out) of the 2019 Copyright Directive.

    Does the four-factor test apply to all media?
    Yes. The 17 USC 107 framework applies regardless of medium, but the application differs.

    How does opt-out interact with fair use?
    Opt-out is not a fair use requirement under US law. It is required under Article 4 of the EU Copyright Directive. As a practical matter, honouring opt-out reduces the fourth factor harm and improves the defence.

    Is non-commercial training automatically fair use?
    No. Non-commercial weighs in the first factor but does not by itself decide the case.

    The four factors applied to AI training

    The fair use analysis under 17 USC 107 considers four factors. Factor one is the purpose and character of the use, including whether the use is of a commercial nature or is for non-profit educational purposes. Factor two is the nature of the copyrighted work. Factor three is the amount and substantiality of the portion used. Factor four is the effect of the use on the potential market for or value of the copyrighted work.

    For AI training each factor presents distinct questions. Factor one turns on whether the training is transformative. The Supreme Court’s 2023 Warhol Foundation v Goldsmith decision narrowed the transformative use analysis, focusing on whether the secondary use has a purpose distinct from the original. Training a general-purpose language model is plausibly transformative. Training a model intended to compete in the same market as the source work is not.

    Factor two distinguishes published from unpublished work and creative from factual work. Published creative works lean against fair use. Factual or functional work leans for fair use. AI training corpora typically contain both, and the factor cuts both ways depending on the sample.

    Factor three considers the amount used. AI training typically copies the entire work into the training pipeline, although the trained model retains only statistical patterns. Courts have split on how to weigh this factor for training. Some treat the full-work ingestion as weighing against fair use. Others focus on what the model retains and find it weighs neutrally or for fair use.

    Factor four is the most contested in AI training cases. Courts ask whether the trained model substitutes for the source work in the relevant market. If the model can output substantially similar content, the substitution effect is direct. If the output is qualitatively different, the substitution is indirect or absent.

    Memorisation and the verbatim output problem

    A central technical question in AI training fair use is memorisation. A trained model that emits verbatim copies of training data has, in effect, retained the training data in its weights. That retention undermines the transformative use argument under factor one and increases market harm under factor four.

    The 2024 and 2025 research literature documented memorisation rates in large language models. Models trained on duplicated content memorise at higher rates. Models trained on rare content memorise at higher rates. Larger models memorise at higher rates. These findings are operational guidance for training pipelines that want to minimise memorisation.

    Mitigation techniques include de-duplication at the document level, de-duplication at the chunk level, training data filtering for high-risk content, and post-training memorisation evaluation with output filtering. Each technique reduces but does not eliminate memorisation. A scraper feeding a training pipeline should apply at least the document-level de-duplication.

    The relationship between fair use and licensing

    Fair use is a defence, not an entitlement. A scraper that has a licence does not need to argue fair use. A scraper without a licence relies on fair use only if the rights holder objects.

    The 2024-2026 trend is toward more licensing. Major publishers struck deals with major AI labs (News Corp with OpenAI in May 2024, Reddit with Google in February 2024, Stack Overflow with OpenAI in May 2024). Those deals reduce reliance on fair use for the licensed corpora.

    For scrapers the implication is that fair use is becoming the fallback, not the default. The strongest position is to license where possible and to fall back to fair use only for content where licensing is impractical.

    Next steps

    If your team trains models on scraped data, the highest-leverage improvement this quarter is to build the training data manifest. It costs little, it underpins both US discovery and EU AI Act response, and it makes every downstream compliance question easier. For broader policy, head to the DRT compliance hub and pair this with the robots.txt ethics guide.

    This guide is informational, not legal advice.

  • CrewAI for scraping pipelines: complete 2026 guide

    CrewAI for scraping pipelines: complete 2026 guide

    CrewAI scraping pipeline patterns turn the abstract “give the LLM some tools and pray” approach into a structured org chart of specialized agents who each do one job well. By early 2026, CrewAI has reached version 0.86, the Crew DSL has stabilized, and the framework’s role-based architecture maps surprisingly cleanly onto the way real scraping teams actually think about work: someone scouts targets, someone fetches, someone extracts, someone validates.

    This guide builds an end-to-end CrewAI scraping pipeline for monitoring competitor prices across multiple ecommerce sites. We define agents, tasks, tools, the hierarchical process, and the integration points with your proxy pool, your database, and your alerting system. By the end you will have a working pipeline plus the knowledge to scale it past prototype.

    Why CrewAI shines for scraping

    CrewAI’s central abstraction is the Crew, a group of Agents that share Tools and execute Tasks under a Process. The framework forces you to think about division of labor up front, and that constraint produces cleaner pipelines than a single mega-agent.

    For scraping, the role split that works in production is:

    Role Responsibility Typical tools
    Scout Discover URLs to scrape Search APIs, sitemap crawler
    Fetcher Pull HTML for each URL with proxy rotation HTTP client, Playwright
    Extractor Parse structured data from HTML LLM with JSON Schema
    Validator QA the extraction against business rules Pydantic, custom validators
    Reporter Format and dispatch results Database writer, notifier

    This shape mirrors how a competent human team handles scraping. CrewAI lets you express it directly.

    Installing the stack

    pip install crewai==0.86.0 crewai-tools==0.20.0 langchain-openai==0.2.10 \
                playwright==1.49.0 pydantic==2.9.2 httpx==0.27.2
    playwright install chromium
    
    export OPENAI_API_KEY="sk-..."
    

    If you prefer Anthropic:

    pip install langchain-anthropic==0.2.10
    export ANTHROPIC_API_KEY="sk-ant-..."
    

    Defining agents

    Agents are declarative. You give them a role, a goal, a backstory (which is more important than it sounds because the LLM uses it for tone and decision style), and a set of tools.

    from crewai import Agent
    from langchain_openai import ChatOpenAI
    
    llm = ChatOpenAI(model="gpt-4o-mini", temperature=0.1)
    
    scout = Agent(
        role="Web Scout",
        goal="Find every product URL for a competitor's catalog",
        backstory=(
            "Senior research analyst with a decade of experience mapping ecommerce "
            "catalogs. Methodical, exhaustive, never misses a category page."
        ),
        llm=llm,
        verbose=True,
        allow_delegation=False,
    )
    
    fetcher = Agent(
        role="HTTP Fetcher",
        goal="Reliably fetch page HTML through a rotating proxy pool",
        backstory=(
            "Pragmatic engineer who treats every fetch as adversarial. Always retries, "
            "always rotates IPs, always respects rate limits."
        ),
        llm=llm,
        verbose=True,
        allow_delegation=False,
    )
    
    extractor = Agent(
        role="Data Extractor",
        goal="Pull title, price, currency, and stock status from product HTML",
        backstory=(
            "Detail-obsessed analyst who treats malformed JSON as a personal insult. "
            "Returns clean structured records or explicit nulls, never guesses."
        ),
        llm=llm,
        verbose=True,
        allow_delegation=False,
    )
    

    Notice how the backstory does most of the work. CrewAI agents read the backstory before every task, and a good backstory shifts behavior more reliably than instruction tweaks in the task itself.

    Building tools

    CrewAI tools are simple Python functions decorated with @tool or subclasses of BaseTool. The tool docstring is what the LLM reads to decide when to call it, so write it like a help message.

    from crewai.tools import BaseTool
    from pydantic import BaseModel, Field
    import httpx
    import os
    import random
    
    PROXIES = os.environ.get("PROXY_POOL", "").split(",")
    
    class FetchInput(BaseModel):
        url: str = Field(..., description="The URL to fetch")
        timeout_s: int = Field(30, description="Request timeout in seconds")
    
    class FetchTool(BaseTool):
        name: str = "fetch_url"
        description: str = (
            "Fetch a URL through the rotating proxy pool. Returns HTML or an error. "
            "Use for any HTTP fetch in this pipeline."
        )
        args_schema: type[BaseModel] = FetchInput
    
        def _run(self, url: str, timeout_s: int = 30) -> str:
            proxy = random.choice(PROXIES) if PROXIES and PROXIES != [""] else None
            with httpx.Client(proxy=proxy, timeout=timeout_s, follow_redirects=True) as c:
                r = c.get(url, headers={"User-Agent": "Mozilla/5.0"})
                return f"HTTP {r.status_code}\n\n{r.text[:200000]}"
    
    fetch_tool = FetchTool()
    

    For Playwright-based fetching of JavaScript-heavy sites, expose a parallel render_url tool:

    from playwright.sync_api import sync_playwright
    
    class RenderTool(BaseTool):
        name: str = "render_url"
        description: str = (
            "Render a URL in headless Chromium and return the HTML after JS executes. "
            "Use only when fetch_url returns insufficient data because of JS-driven content."
        )
        args_schema: type[BaseModel] = FetchInput
    
        def _run(self, url: str, timeout_s: int = 30) -> str:
            with sync_playwright() as p:
                browser = p.chromium.launch(headless=True)
                page = browser.new_page()
                page.goto(url, wait_until="networkidle", timeout=timeout_s * 1000)
                html = page.content()
                browser.close()
            return html
    
    render_tool = RenderTool()
    

    Attach the tools to the right agents:

    fetcher.tools = [fetch_tool, render_tool]
    

    Defining tasks

    Tasks express what each agent should do and what they should produce. Every task gets an expected_output description that doubles as the validation hint for the LLM.

    from crewai import Task
    
    scout_task = Task(
        description=(
            "Discover all product URLs for {target_site} in the {category} category. "
            "Use the sitemap if available. Return a JSON array of URLs."
        ),
        expected_output="JSON array of strings, each a fully qualified product URL",
        agent=scout,
    )
    
    fetch_task = Task(
        description=(
            "For each URL in the prior task output, fetch the HTML using fetch_url. "
            "If the response contains less than 1000 characters of body content or "
            "looks like a JS-only shell, retry with render_url. "
            "Return a JSON map of {url: html}."
        ),
        expected_output="JSON object mapping URL to raw HTML",
        agent=fetcher,
        context=[scout_task],
    )
    
    extract_task = Task(
        description=(
            "For each {url: html} pair, extract title, price (number), currency (3-letter code), "
            "and in_stock (boolean). Return a JSON array of records."
        ),
        expected_output=(
            "JSON array of objects with keys: url, title, price, currency, in_stock"
        ),
        agent=extractor,
        context=[fetch_task],
    )
    

    The context field is how data flows between tasks. CrewAI passes prior task outputs as text to the next agent’s prompt.

    Assembling the crew

    from crewai import Crew, Process
    
    crew = Crew(
        agents=[scout, fetcher, extractor],
        tasks=[scout_task, fetch_task, extract_task],
        process=Process.sequential,
        verbose=True,
    )
    
    result = crew.kickoff(inputs={
        "target_site": "https://www.lazada.sg",
        "category": "ergonomic-keyboards",
    })
    
    print(result.raw)
    

    Sequential is the default and the right pick for a linear scraping pipeline. For more complex workflows where one agent should orchestrate others, use Process.hierarchical and provide a manager LLM.

    Hierarchical process for adaptive pipelines

    The hierarchical process puts a manager agent in charge. The manager decides which worker to call, in what order, with what arguments. This is the right shape when the pipeline shape depends on the input.

    from crewai import Crew, Process
    from langchain_openai import ChatOpenAI
    
    manager_llm = ChatOpenAI(model="gpt-4o", temperature=0)
    
    crew = Crew(
        agents=[scout, fetcher, extractor],
        tasks=[scout_task, fetch_task, extract_task],
        process=Process.hierarchical,
        manager_llm=manager_llm,
        verbose=True,
    )
    

    The manager dispatches dynamically and is the right choice when you cannot enumerate the steps up front. Cost is higher because every dispatch decision is an LLM call. Reserve for genuine adaptivity, not as the default.

    Comparing CrewAI to LangGraph and AutoGen

    Dimension CrewAI LangGraph AutoGen
    Mental model Org chart of roles State machine Group chat
    Best fit for scraping Multi-step pipelines with clear roles Branching state-driven workflows Conversational extraction
    Tool definition BaseTool subclass LangChain Tool Function decorator
    Persistence Manual Built-in checkpointer Manual
    Learning curve Lowest Moderate Steep
    Production maturity High High Moderate
    Native MCP Via wrappers Via wrappers Yes

    CrewAI is the framework to pick when the pipeline shape mirrors a team org chart and you want the fastest possible path from idea to working code. LangGraph wins for pipelines with non-linear branching and long-running state. AutoGen wins for conversational extraction where multiple agents debate the right answer.

    For the LangGraph alternative, see our scraping with LangGraph agents guide. For AutoGen, see Multi-agent scraping with AutoGen in 2026.

    Adding proxy rotation

    CrewAI itself does not own the network layer; your tools do. The pattern is to keep a proxy pool in env or in a small singleton, and pick from it inside every fetch tool. We showed this pattern above; here is the production refinement that adds health checking.

    import time
    from collections import defaultdict
    
    class ProxyPool:
        def __init__(self, proxies):
            self.proxies = proxies
            self.failures = defaultdict(int)
            self.last_used = defaultdict(float)
    
        def pick(self):
            now = time.time()
            candidates = [p for p in self.proxies if self.failures[p] < 3]
            if not candidates:
                self.failures.clear()
                candidates = self.proxies
            return min(candidates, key=lambda p: self.last_used[p])
    
        def report(self, proxy, success):
            self.last_used[proxy] = time.time()
            if not success:
                self.failures[proxy] += 1
            else:
                self.failures[proxy] = max(0, self.failures[proxy] - 1)
    
    pool = ProxyPool([os.environ["PROXY_POOL"].split(",")])
    

    For ASEAN ecommerce scraping with mobile IPs that pass strict carrier-level checks, Singapore mobile proxy plugs into this pool directly.

    Validator and Reporter agents in detail

    The five-role split listed at the top of the article needs two agents we have not yet shown in code. Adding them turns a fragile demo into a production-grade pipeline.

    from crewai import Agent
    
    validator = Agent(
        role="Data Quality Auditor",
        goal=(
            "Reject any extracted record that violates business rules. Price must be > 0 "
            "and < 100000. Currency must be a valid ISO 4217 code. Title must be non-empty "
            "and under 500 chars. in_stock must be a bool. Flag suspicious values for review."
        ),
        backstory=(
            "Former data engineer who spent two years cleaning a B2B product catalog. "
            "Allergic to silently incorrect data. Will refuse to pass through anything "
            "that smells wrong, and will document why."
        ),
        llm=llm,
        verbose=True,
    )
    
    reporter = Agent(
        role="Insights Reporter",
        goal=(
            "Compare today's extracted prices against yesterday's and emit a Slack-ready "
            "summary highlighting price drops over 10%, new SKUs, and out-of-stock changes."
        ),
        backstory=(
            "Pricing analyst with a journalism background. Writes summaries that a busy "
            "merchandising manager can act on in 30 seconds."
        ),
        llm=llm,
        verbose=True,
    )
    

    The Validator agent in particular pays for itself the first time the Extractor mistakes a postcode field for a price and tries to write 94025 to your price_usd column.

    Custom Python validation tool

    LLMs are fine for fuzzy validation but bad at strict rule checking. Pair the Validator agent with a deterministic tool that does the unforgiving work.

    from crewai.tools import BaseTool
    from pydantic import BaseModel, ValidationError, Field as PField, conlist
    from typing import List, Literal
    
    class ProductRecord(BaseModel):
        url: str
        title: str = PField(min_length=1, max_length=500)
        price: float = PField(gt=0, lt=100000)
        currency: Literal["USD", "SGD", "EUR", "JPY", "GBP", "INR", "MYR", "THB", "IDR", "VND"]
        in_stock: bool
    
    class ValidateInput(BaseModel):
        records: list
    
    class StrictValidatorTool(BaseTool):
        name: str = "strict_validate"
        description: str = (
            "Run deterministic validation on a list of product records. Returns the "
            "subset that passes plus a list of errors for the rejected records."
        )
        args_schema: type[BaseModel] = ValidateInput
    
        def _run(self, records: list) -> dict:
            passed, errors = [], []
            for r in records:
                try:
                    passed.append(ProductRecord(**r).model_dump())
                except ValidationError as e:
                    errors.append({"record": r, "errors": e.errors()})
            return {"passed": passed, "errors": errors}
    

    The Validator agent calls strict_validate and uses its output to decide what to forward to the Reporter.

    Adding async execution and concurrency

    Sequential is the default but real production crews need parallelism. CrewAI 0.86 introduced Process.async_sequential and async-friendly task callbacks. Use them when tasks are independent.

    from crewai import Crew, Process
    
    crew = Crew(
        agents=[scout, fetcher, extractor, validator, reporter],
        tasks=[scout_task, fetch_task, extract_task, validate_task, report_task],
        process=Process.sequential,
    )
    
    # Process N URLs in parallel by spawning N crews
    import asyncio
    
    async def scrape_many(urls):
        sem = asyncio.Semaphore(10)  # cap concurrent crews
    
        async def one(url):
            async with sem:
                return await crew.kickoff_async(inputs={"target_url": url})
    
        return await asyncio.gather(*(one(u) for u in urls))
    

    The Semaphore is the single most important line. Without it, 1000 URLs spawn 1000 simultaneous Crew instances, each holding a Playwright browser, and the host runs out of memory in 30 seconds.

    Wiring CrewAI into a warehouse

    The Reporter agent should not be writing directly to your warehouse. Give it a thin tool that calls a typed function in your data layer.

    import asyncpg
    
    class WarehouseWriteTool(BaseTool):
        name: str = "warehouse_write"
        description: str = "Append validated product records to the warehouse."
    
        def _run(self, records: list) -> dict:
            # synchronous wrapper around an async pool
            return {"written": _sync_warehouse_write(records)}
    

    The reason to keep this thin: if the LLM hallucinates a malformed record that slips through the Validator, the warehouse layer catches it via the column-type contract. Defense in depth.

    Cost benchmarks

    For a sequential scrape of 100 product URLs with three agents:

    Setup LLM cost Wall clock
    GPT-4o-mini all agents $0.45 8 min
    GPT-4o all agents $7.20 9 min
    Mixed: 4o-mini scout/fetcher, 4o extractor $1.80 8.5 min
    Claude Haiku all agents $0.55 7 min
    Claude Sonnet all agents $7.80 8 min

    The mixed setup is the value sweet spot. Use cheap models for orchestration agents, expensive models only for the agent that needs to read messy HTML and produce clean JSON.

    Cost levers in priority order

    The fastest wins on CrewAI cost, in the order they pay off:

    1. Trim HTML before passing to Extractor. Strip scripts, styles, comments. Cuts tokens by 40 to 70 percent.
    2. Use GPT-4o-mini for Scout, Fetcher, Validator, Reporter. Reserve GPT-4o or Sonnet for Extractor only.
    3. Cache extractions by HTML hash. Same page extracted twice should not pay LLM twice.
    4. Cap max_iter per agent to 8. Default 25 lets confused agents loop expensively.
    5. Switch verbose off in production. Verbose mode includes all intermediate thought tokens in the trace which the agent re-reads.

    Combined, these cut typical per-page cost by 60 to 80 percent on a tuned pipeline versus a default-config baseline.

    Memory and learning

    CrewAI agents can be configured with memory that persists across runs. Two flavors: short-term (within a crew run) and long-term (across runs, backed by a vector store).

    from crewai.memory import LongTermMemory
    from crewai.memory.storage.ltm_sqlite_storage import LTMSQLiteStorage
    
    extractor.memory = True
    
    crew = Crew(
        agents=[scout, fetcher, extractor],
        tasks=[scout_task, fetch_task, extract_task],
        process=Process.sequential,
        memory=True,
        long_term_memory=LongTermMemory(
            storage=LTMSQLiteStorage(db_path="./crew_ltm.db")
        ),
    )
    

    For scraping, long-term memory pays off when the same agent learns site-specific quirks across runs. After a few iterations, the extractor remembers that a certain Lazada page renders price in a non-standard div.

    Production deployment

    Run CrewAI under a worker queue (Celery, RQ, or a Postgres-backed queue) and treat each crew kickoff as a job. Set hard timeouts, log every agent step, and persist the result to your warehouse.

    For long-running crews, set step_callback and task_callback to stream progress to your observability stack. The callbacks fire after every agent step and every task completion respectively.

    The official CrewAI documentation covers deployment options in depth, including the managed CrewAI Plus platform.

    Frequently asked questions

    Can CrewAI agents call MCP servers?
    Not natively in 0.86. Wrap the MCP server in a Python tool that translates BaseTool calls to MCP JSON-RPC. The community has at least three open-source bridges; pick the one most actively maintained.

    How do I prevent runaway agent loops?
    Set max_iter on each agent (default 25). For task-level safety, set max_execution_time in seconds. Both options cap cost and clock.

    Can I run CrewAI without OpenAI or Anthropic?
    Yes. Any LangChain LLM works, including Ollama, vLLM, LM Studio, and Bedrock. Quality drops with smaller open-source models, especially on the extraction agent which needs strong JSON Schema adherence.

    Does CrewAI handle parallel agent execution?
    Sequential and hierarchical processes are single-threaded. For genuine parallelism, run multiple crew invocations under asyncio or a worker pool.

    What about debugging when an agent goes off-script?
    Set verbose=True on agents and crew. The terminal trace shows every thought, action, and observation. For richer logging, integrate with LangSmith.

    Are CrewAI tasks idempotent? Can I safely retry?
    Tasks themselves are idempotent only if your tools are idempotent. The framework will not deduplicate side effects. Wrap the side-effecting step (warehouse write, alert dispatch) with an idempotency key based on the input.

    Can the same crew handle multiple sites with different layouts?
    Yes, but the Extractor agent benefits from per-site backstory or per-site fewshot examples. The pattern that works is one Crew per site, sharing the Scout/Fetcher/Validator/Reporter agents and swapping only the Extractor.

    How does CrewAI compare to writing the same pipeline in plain LangChain?
    CrewAI is roughly 30 percent fewer lines of code for typical 3-to-5 agent pipelines and the role-based abstraction makes the intent clearer in code review. Plain LangChain wins when you need fine-grained control over the prompt assembly per turn.

    Can I use CrewAI for crawling, not just scraping?
    Yes. Add a Crawler agent with a BFS tool that takes seed URLs and depth, and pass the discovered URL list to the Fetcher. CrewAI’s role abstraction handles the recursion via the same Crawler agent re-invoked per depth.

    Common production gotchas

    A handful of issues bite teams during their first month.

    The Extractor agent silently truncates HTML when the input exceeds the LLM’s context window. Always pre-trim with a deterministic rule (strip scripts, styles, hidden divs, and whitespace) before sending to the LLM, and log the trimmed size.

    CrewAI’s verbose=True writes to stdout in a hard-to-parse format. For production, replace with a structured logging callback that emits JSON lines you can grep and ship to your log aggregator.

    Tool argument schemas are inferred at class load time. Adding a field after the agent has already been constructed is a no-op until you reload. Restart workers on tool changes.

    Hierarchical mode’s manager LLM is GPT-4o by default and that single decision dominates the cost on small jobs. Override manager_llm to GPT-4o-mini or a stronger model only when the dispatch decision is hard.

    The crew’s kickoff method blocks the calling thread. For async workers, use kickoff_async and await it. Mixing the two in the same process leads to nested event loop errors.

    For broader patterns on building agentic scraping in 2026, browse the AI agentic proxies category.

  • The HiQ vs LinkedIn ruling: what scrapers should know in 2026

    The HiQ vs LinkedIn ruling: what scrapers should know in 2026

    The HiQ LinkedIn scraping ruling is one of the most cited court decisions in the entire web data space, and one of the most misunderstood. Almost every commercial scraper today operates under a mental model of “scraping public data is legal because of HiQ.” That mental model is partially correct, partially wrong, and partially incomplete in 2026 in ways that matter for how you build your pipeline. This guide walks through what the case actually decided, how Van Buren v. United States changed the landscape in 2021, what happened after the case eventually settled in 2022, and what the practical takeaway is for scrapers operating in 2026.

    The audience is the technical lead, in-house counsel, or product owner who has heard the case name dropped in vendor pitches and customer conversations and wants the actual story.

    What the case was actually about

    HiQ Labs was a small data analytics company that scraped publicly visible LinkedIn profiles to build a workforce analytics product. They sold predictions about employee flight risk to large enterprise customers. LinkedIn sent HiQ a cease-and-desist letter in 2017 demanding they stop. HiQ sued LinkedIn for a declaratory judgement that their scraping was lawful and for an injunction preventing LinkedIn from blocking them.

    The dispute was framed primarily under the Computer Fraud and Abuse Act (CFAA), 18 U.S.C. Section 1030. LinkedIn argued that HiQ’s continued scraping after the cease-and-desist letter constituted “access without authorisation” under the CFAA. HiQ argued that publicly visible data, accessible without a login, could not be “without authorisation” because no authorisation was required in the first place.

    The Northern District of California granted HiQ a preliminary injunction in 2017. The Ninth Circuit affirmed in 2019. The Supreme Court vacated and remanded in 2021 in light of Van Buren. The Ninth Circuit affirmed again on remand in 2022. The case eventually settled later that year, with HiQ agreeing to certain limits and LinkedIn dropping its claims.

    For the broader compliance picture across jurisdictions, see the GDPR compliance guide and the CCPA compliance guide.

    What the Ninth Circuit actually held

    The Ninth Circuit in 2019 (and again in 2022) held that the CFAA’s “without authorisation” language likely does not cover access to publicly available websites, because for such sites no authorisation is needed in the first place. The court drew an analogy: a cease-and-desist letter does not transform a public sidewalk into private property. The walking is the access; the public availability is the authorisation.

    The court emphasised three points:

    1. The CFAA was originally enacted to address computer hacking, not civil disputes about access to public webpages.
    2. Reading “without authorisation” broadly to include cease-and-desist-revocation would create a “criminal” regime in which any TOS violation became a federal crime.
    3. The First Amendment and antitrust concerns weighed against allowing platforms to use the CFAA as a private weapon to control access to public information.

    The decision did not say all scraping is legal. It said the CFAA does not criminalise scraping of publicly accessible data merely because the platform sent a cease-and-desist letter.

    How Van Buren changed the landscape in 2021

    In June 2021, the Supreme Court decided Van Buren v. United States. The case did not involve scraping, but it interpreted the CFAA’s “exceeds authorised access” language. A police officer had used his lawful access to a database to look up information for an improper purpose. The government argued he had “exceeded” his authorisation because his access was conditioned on use for proper purposes.

    The Supreme Court rejected this reading. It held that “exceeds authorised access” means accessing files, folders, or databases that are off-limits, not accessing permitted data for impermissible purposes. The “gates-up-or-down” model: if the gate is up, your access is not unauthorised even if the use is.

    This narrow reading of the CFAA aligned exactly with the Ninth Circuit’s HiQ analysis, and the Supreme Court remanded HiQ for reconsideration in light of Van Buren. The Ninth Circuit, on remand, reaffirmed its earlier decision.

    The combined effect: in the Ninth Circuit (covering California and most of the western US), the CFAA does not reach scraping of publicly available websites, regardless of TOS violations or cease-and-desist letters.

    What the eventual 2022 settlement actually said

    Most coverage of HiQ stopped at the Ninth Circuit affirmance. Less covered: the case settled in late 2022. The settlement terms were partly confidential, but several public components were disclosed. HiQ agreed to a permanent injunction prohibiting it from scraping LinkedIn member data going forward. LinkedIn dropped its remaining claims.

    Why did HiQ agree to a permanent injunction if they had won? Two reasons. First, after years of litigation, the company had effectively lost commercial momentum and was wound down. Second, while the CFAA claim had been knocked out, LinkedIn’s parallel state-law claims (breach of contract for TOS violations, trespass to chattels, tortious interference) were still live. The CFAA win did not resolve the state-law theories.

    This is the part most scrapers miss. Winning the CFAA fight does not win the state-law fight. Public availability does not erase contract. A scraper that creates an account, agrees to terms, and then scrapes is in a fundamentally different posture than a scraper that hits public URLs without ever logging in.

    The Meta v. Bright Data parallel

    In January 2024, the US District Court for the Northern District of California decided Meta v. Bright Data. The fact pattern was deliberately similar to HiQ: Bright Data scraped public Facebook and Instagram pages and resold the data. Meta sued for breach of contract (TOS violation) and tortious interference.

    The court held that Meta’s TOS only bound logged-in users. Bright Data’s scraping of publicly accessible logged-out pages did not breach the contract because no contract had ever formed. This was a very strong scraper-side ruling, but it was narrow: it applied only to logged-out scraping. The moment a scraper authenticates, the TOS attaches.

    The Israeli court ruled similarly in a parallel proceeding the same year. EU regulators were quick to point out that the absence of a contract violation does not equal a lawful basis under GDPR, but the court rulings did establish a clear US-side rule: logged-out scraping of public data is legally robust against TOS-based claims.

    For the broader public-vs-personal data analysis, see the personal vs public data scraping framework.

    What the rulings collectively permit and forbid

    Activity Legal posture in 2026 (US)
    Scraping public URLs, no login Generally permitted; CFAA does not reach
    Scraping behind a login you created TOS applies; breach-of-contract risk
    Scraping after a cease-and-desist letter (logged out) Permitted under HiQ
    Scraping behind a paywall you bypassed High risk; CFAA may reach (gate is down)
    Bulk personal data resale TOS-independent risks: state privacy law
    Scraping for AI training Legal under HiQ; copyright fair use is separate
    Scraping with fake accounts TOS breach; state-law exposure

    The asymmetry between logged-in and logged-out is the most important practical takeaway. Public, logged-out scraping has a strong legal floor in the US. Anything behind authentication exists under the platform’s terms.

    Decision tree: is your scraping covered by HiQ?

    Q1: Is the target URL accessible without any login?
        ├── No  -> HiQ does not protect you. Evaluate TOS and CFAA.
        └── Yes -> Q2
    Q2: Are you in the Ninth Circuit's jurisdiction (or a court likely to follow)?
        ├── Yes -> CFAA risk is low.
        └── No  -> Other circuits have not all adopted; evaluate locally.
    Q3: Did you create an account and accept terms?
        ├── Yes -> TOS attaches; CFAA may not, but contract claim does.
        └── No  -> Q4
    Q4: Are you scraping personal data of identifiable individuals?
        ├── Yes -> CCPA, GDPR, PDPA, DPDP may apply independently.
        └── No  -> Q5
    Q5: Are you bypassing technical access controls (CAPTCHA, IP block, rate limit)?
        ├── Yes -> CFAA "gate is down" risk increases.
        └── No  -> Strongest defensive posture.
    

    The combination of “logged-out, public, identifiable-but-public, no controls bypassed” is the strongest position. Each “yes” to authentication, controls bypass, or personal data adds risk that HiQ does not resolve.

    Practical implications for 2026 scraping pipelines

    Three operational implications.

    First, separate your logged-out and logged-in scraping infrastructure. Logged-out scraping enjoys HiQ-grade protection. Logged-in scraping operates under TOS and contract law. Mixing the two in one pipeline obscures your legal posture and weakens both.

    Second, do not bypass technical access controls. The Van Buren “gates-up-or-down” model means that if a site puts up a gate (CAPTCHA, IP block, paywall) and you bypass it, you have moved from “authorisation not required” to “authorisation explicitly denied.” That is a different legal universe.

    Third, document your access methodology. If you ever need to invoke HiQ in your defence, you will need to prove that your access was logged-out, that no controls were bypassed, and that the data was genuinely publicly accessible. A scrape that goes through a residential proxy mesh after using browser fingerprint spoofing to defeat a fingerprinting check is not “publicly accessible” in the HiQ sense.

    For a deeper dive on the bot management and fingerprinting question, see the DataDome vs PerimeterX vs Akamai comparison.

    External references

    The Ninth Circuit opinion in HiQ Labs v LinkedIn (2022 remand) is at cdn.ca9.uscourts.gov/datastore/opinions/2022/04/18/17-16783.pdf. The Supreme Court opinion in Van Buren v United States is at supremecourt.gov/opinions/20pdf/19-783_k53l.pdf. The Meta v Bright Data summary judgement is in the public PACER record for case 3:23-cv-00077.

    Comparison: HiQ doctrine vs state contract law vs GDPR

    Issue HiQ / Van Buren (CFAA) State contract law GDPR
    Reaches public URLs No No (no contract) Yes
    Reaches logged-in scraping Limited Yes (TOS) Yes
    Requires lawful basis No No Yes
    Personal data carve-out Irrelevant Irrelevant Public availability not a defence
    Statutory damages Up to USD 1k per access Variable Up to 4% revenue
    Cease-and-desist effect None on logged-out Strengthens contract claim Irrelevant
    Settlement strategy lever Weak (clean win) Strong (TOS hook) Strong (DPA leverage)

    The takeaway: HiQ is a strong shield against CFAA claims for logged-out public scraping. It is not a shield against contract, GDPR, or other state privacy law claims. Build your legal posture for all four regimes simultaneously.

    What changed in 2024-2025 case law

    Three additional cases shaped the 2026 landscape.

    X Corp v Bright Data (2024, ND Cal) reaffirmed Meta v Bright Data: logged-out scraping is not a TOS breach because no contract attaches. The court was explicit that platform terms cannot bind non-users.

    Reddit v Anthropic (2025, ND Cal) is still pending as of mid-2026. Reddit alleges Anthropic scraped past explicit robots.txt directives blocking ClaudeBot. The case will test whether ignoring robots.txt for AI training constitutes any kind of cognisable claim distinct from CFAA. The outcome will reshape AI scraping practice.

    Doe v Github (2024, ND Cal) addressed scraping of open-source code for AI training, holding that the public availability of code on GitHub did not waive copyright protections in derivative or memorised outputs. Public availability and copyright protection are separate inquiries.

    For a forward-looking discussion of where the AI training case law is heading, see fair use for AI training data in 2026.

    FAQ

    Is all web scraping legal because of HiQ?
    No. HiQ knocked out one specific federal claim (the CFAA) for one specific kind of scraping (logged-out public data). It did not legalise all scraping. Contract, copyright, privacy, and trade-secret claims survive independently.

    Does HiQ apply outside the Ninth Circuit?
    The Ninth Circuit covers California and the western US, where most tech litigation lands. Other circuits have not uniformly adopted the same reading, but the trend post-Van Buren is in that direction.

    What about scraping social media platforms?
    Logged-out, public-page scraping is reasonably defensible under HiQ and Meta v Bright Data. Logged-in scraping is a TOS issue and should be evaluated separately.

    Can a cease-and-desist letter make my scraping illegal?
    Under the CFAA in the Ninth Circuit, no. Under state contract or trespass law, it can strengthen the platform’s claim. Treat a C&D as a serious signal even if it does not change the federal analysis.

    Did HiQ actually win in the end?
    The CFAA fight, yes. The commercial fight, no. HiQ wound down operations and agreed to a permanent injunction in the 2022 settlement. The case is a legal win and a commercial cautionary tale.

    Extended case law analysis

    The hiQ Labs v LinkedIn litigation ran from 2017 to 2022 and produced four major opinions. The 2019 Ninth Circuit opinion held that scraping public data did not violate the CFAA’s without authorisation prong because public data is by definition authorised for any visitor. The Supreme Court’s Van Buren v United States decision (June 2021) reinforced the gates-up gates-down reading of the CFAA, which strongly supported the hiQ position. The 2022 Ninth Circuit opinion on remand reaffirmed the public-data holding and remanded the contract claims, which hiQ ultimately settled.

    The post-hiQ landscape contains four important precedents that scrapers should know.

    1. Meta Platforms v Bright Data (Northern District of California, January 2024). Meta’s CFAA and contract claims against Bright Data largely failed at summary judgment for public data scraping. The court relied on hiQ for the CFAA analysis and held that Bright Data had not formed a contract by browsing logged-out pages.

    2. X Corp v Bright Data (Northern District of California, May 2024). The court reached a similar conclusion, dismissing X’s claims for scraping public data while users were not logged in.

    3. Ryanair v PR Aviation (CJEU, 2015) and the 2024 follow-ups. The European pathway places more weight on database rights and contract than on CFAA-equivalent statutes.

    4. The 2024-2025 wave of state-level scraping statutes, including bills introduced in California, Texas, and New York that propose explicit rules for AI training data scraping.

    Implementation patterns post-hiQ

    Operators acting on the hiQ ruling should follow a five-step posture.

    1. Distinguish logged-out from logged-in scraping. The hiQ holding extends only to logged-out public scraping. Logged-in scraping involves account terms which hiQ does not protect.

    2. Avoid technical bypass of authentication, rate limits, or bot detection beyond ordinary headless browser use. Bypass is what triggers the CFAA in post-hiQ cases.

    3. Maintain a documented record of the scrape’s purpose, frequency, and destination. Litigation discovery will surface this and a clean record helps.

    4. Honour cease and desist letters carefully. The hiQ ruling does not give scrapers a right to ignore valid C and D letters that allege contract or tort claims. The right response is legal review, not silence.

    5. Apply privacy law independently. CFAA protection does not equal GDPR or CCPA protection. Personal data is regulated by separate statutes.

    Code pattern: distinguishing logged-out scraping

    def is_logged_out_only(session):
        if session.cookies:
            return False
        if "Authorization" in session.headers:
            return False
        if "X-Auth-Token" in session.headers:
            return False
        return True
    

    Comparison: scraping legal posture before and after hiQ

    Question Pre-hiQ default Post-hiQ default (US public data) Caveat
    CFAA risk for public scraping Substantial Low Bypass changes the analysis
    Contract risk Moderate Moderate to high Browsewrap terms still litigated
    Trespass to chattels Low Low Resource exhaustion claims survive
    Privacy law risk High High Independent of CFAA
    Copyright Moderate Moderate Fair use defence is fact-specific

    Additional FAQ

    Does hiQ apply to data behind a login?
    No. The holding is limited to public, unauthenticated data. Logged-in scraping involves account terms.

    Does hiQ apply outside the United States?
    No. Each jurisdiction has its own statutes. The European Union, the United Kingdom, Singapore, and India apply different frameworks.

    Can a publisher block scrapers technically?
    Yes. Technical blocks (rate limits, IP bans, bot detection) are lawful. Bypassing them weakens the hiQ defence.

    Is the hiQ ruling settled law?
    The Ninth Circuit holding stands. Other circuits have not directly contradicted it but they could. A scraping operation should not assume nationwide uniformity.

    Cases that built on the hiQ framework

    Two post-hiQ decisions have shaped how courts apply the doctrine in 2024-2026.

    In Meta Platforms v. Bright Data (N.D. Cal., January 2024), Judge Edward Chen granted summary judgment to Bright Data on Meta’s contract claims for scraping logged-out Facebook and Instagram public profile data. The court explicitly relied on the hiQ framework, holding that Meta could not enforce its terms of service against Bright Data because Bright Data had no account and had not assented to the terms. The decision reinforced the bright-line significance of the logged-in versus logged-out distinction. Meta’s separate trespass and unjust enrichment theories were also dismissed.

    In X Corp v. Bright Data (N.D. Cal., May 2024), Judge William Alsup followed the same logic for logged-out scraping of X (formerly Twitter) public posts. Judge Alsup’s opinion went further than Meta v. Bright Data in characterising the policy stakes, observing that giving social media platforms unilateral power to control public-facing data would create de facto information monopolies inconsistent with US antitrust and free-speech traditions. The decision is now the most quotable post-hiQ pro-scraping precedent for logged-out commercial use.

    Both cases stop at logged-out scraping. Neither protects the bypass of authentication, the use of fake accounts, or the circumvention of technical blocks. The combined doctrine is narrow but settled in the Ninth Circuit for the use cases it covers.

    The litigation history of hiQ

    The hiQ Labs v LinkedIn dispute began in May 2017 when LinkedIn sent hiQ a cease and desist letter demanding that hiQ stop scraping public LinkedIn profiles. hiQ sued for declaratory relief, arguing that scraping public data did not violate the Computer Fraud and Abuse Act and that LinkedIn’s blocking efforts violated antitrust and tortious interference principles.

    The Northern District of California granted hiQ a preliminary injunction in August 2017. The Ninth Circuit affirmed in September 2019, holding that scraping public data did not violate the CFAA’s without authorisation prong. The Supreme Court vacated and remanded in light of Van Buren v United States in June 2021. The Ninth Circuit affirmed again on remand in April 2022, reaffirming the public-data holding.

    The case finally settled in late 2022 after the contract claims were remanded to the district court. The settlement terms were not made public, but hiQ agreed to stop scraping LinkedIn and acknowledged having breached LinkedIn’s terms of service. The company subsequently wound down operations.

    The procedural history matters because the holding stands even though the company that brought the case did not survive. The legal precedent is what scrapers operate under in 2026, and that precedent is favourable to public-data scraping but does not provide a shield against contract claims.

    How the Supreme Court’s Van Buren decision changed the analysis

    Van Buren v United States, decided in June 2021, addressed the without authorisation and exceeds authorised access prongs of the CFAA. The Supreme Court adopted a gates-up gates-down reading, under which a person violates the CFAA only when they access a computer system that is closed to them. A person who is authorised to access certain files but uses that access for an improper purpose does not exceed authorised access.

    Applied to scraping, Van Buren strengthened the hiQ holding. Public LinkedIn profiles are gates-up for any visitor. A scraper that accesses them is not exceeding authorised access. The CFAA does not reach the activity.

    Van Buren did not reach the question of when a scraper is gates-down. Subsequent cases have explored that question. A scraper that bypasses authentication, ignores IP blocks, or rotates User-Agent strings to evade detection is plausibly gates-down. The line is fact-specific and remains unsettled.

    Practical operational implications

    A 2026 scraping operation taking hiQ as authoritative should adopt five operational practices.

    First, document the public-facing nature of every target page. Maintain screenshots showing that the page is reachable without login. Maintain timestamps. The evidentiary record matters in litigation discovery.

    Second, avoid technical bypass. The hiQ holding does not protect a scraper that bypasses CAPTCHA, defeats bot detection, or rotates IPs to evade rate limits. Each of those activities is a potential gates-down trigger.

    Third, respond promptly to cease and desist letters. Ignoring a C and D letter does not improve the legal position. The right response is legal review followed by either a measured response or a tactical pause.

    Fourth, separate logged-in scraping from logged-out scraping operationally. Use different infrastructure, different credentials, different audit trails. The legal analysis is fundamentally different.

    Fifth, monitor circuit splits. The hiQ ruling is Ninth Circuit law. Other circuits may reach different conclusions. A scraper operating nationally should track relevant cases in the Second, Fourth, and Eleventh Circuits.

    Next steps

    If your team relies on HiQ in any pitch, customer conversation, or compliance memo, the fastest improvement is to make the logged-out vs logged-in distinction explicit in your documentation. The legal protections are very different. For a fuller compliance posture across regimes, head to the DRT compliance hub and pair this with the GDPR and CCPA guides.

    This guide is informational, not legal advice.

  • Scraping with LangGraph agents in 2026

    Scraping with LangGraph agents in 2026

    LangGraph scraping agents have become the standard pattern for any non-trivial LLM-driven scraping pipeline that needs branching, retries, and checkpointed state. By early 2026, LangGraph has reached version 0.4 with stable APIs, the StateGraph primitive is rock solid, and the Postgres checkpointer makes it easy to resume long-running scrapes after a crash. If your scraping job is more than fetch-extract-store, LangGraph is the right framework.

    This guide builds a real production scraping agent step by step. We define the state, wire the tool nodes, add retry edges, plug in a checkpointer, and benchmark cost against alternatives. By the end you will have a working LangGraph scraper that handles a flaky target site, recovers from failures, and emits clean structured data.

    Why LangGraph beats LangChain agents for scraping

    LangGraph scraping agents express the workflow as an explicit graph instead of an implicit ReAct loop. That difference matters in three places.

    First, branching. A scraping agent often needs to take different paths depending on what the page returns. Did the site return a captcha? Branch to the solver. Is the price hidden behind a login? Branch to authentication. LangChain’s ReAct agent makes branching implicit through prompt engineering. LangGraph makes it explicit through edges.

    Second, observability. When a scraper fails at 3 AM, you want to know exactly which node failed and what state was in scope. LangGraph’s state object plus LangSmith integration gives you that. The classic LangChain agent gives you a chain of thought that you have to read.

    Third, persistence. LangGraph ships a checkpointer system. After every node, the state is persisted to SQLite or Postgres. If the worker dies mid-scrape, you resume from the last checkpoint with one line of code.

    Where LangGraph fits next to LangChain

    LangChain remains the right framework for prompt templates, LLM clients, retrievers, and document loaders. LangGraph builds on top, adding the runtime that orchestrates them. The mental model is: LangChain is the parts bin, LangGraph is the assembly line. Almost every production LangGraph scraping agent imports LangChain primitives for the LLM call and the prompt template, and reserves LangGraph for the routing and state machine.

    Installing the stack

    pip install langgraph==0.4.0 langchain==0.3.20 langchain-openai==0.2.10 \
                langgraph-checkpoint-postgres==2.0.10 \
                playwright==1.49.0 pydantic==2.9.2 httpx==0.27.2
    playwright install chromium
    

    For LangSmith tracing (free for personal projects):

    export LANGSMITH_API_KEY="ls__..."
    export LANGSMITH_TRACING="true"
    export LANGSMITH_PROJECT="lazada-scraper"
    

    Defining the agent state

    The state is a Pydantic-style TypedDict that flows through every node. For a scraping agent it typically holds the input URL, the fetched HTML, the extracted data, and an error log.

    from typing import TypedDict, Optional, List
    from langgraph.graph import StateGraph, END
    
    class ScrapeState(TypedDict):
        url: str
        html: Optional[str]
        captcha_detected: bool
        extracted: Optional[dict]
        errors: List[str]
        attempt: int
    

    LangGraph reduces state by merging dicts on every node return. Mutations to the state inside a node are not seen by other nodes; only the returned dict is.

    Building the nodes

    Each node is a function that takes the state and returns a partial state update.

    import asyncio
    from playwright.async_api import async_playwright
    from openai import AsyncOpenAI
    import json
    
    client = AsyncOpenAI()
    
    async def fetch_node(state: ScrapeState) -> dict:
        """Fetch the URL with Playwright."""
        try:
            async with async_playwright() as p:
                browser = await p.chromium.launch(headless=True)
                page = await browser.new_page()
                await page.goto(state["url"], wait_until="networkidle", timeout=30000)
                html = await page.content()
                await browser.close()
            return {"html": html, "attempt": state["attempt"] + 1}
        except Exception as e:
            return {
                "errors": state["errors"] + [f"fetch failed: {e}"],
                "attempt": state["attempt"] + 1,
            }
    
    async def captcha_check_node(state: ScrapeState) -> dict:
        """Quick heuristic for captcha detection."""
        html = state.get("html", "") or ""
        flags = ["cf-challenge", "captcha", "px-captcha", "datadome", "recaptcha"]
        detected = any(f in html.lower() for f in flags)
        return {"captcha_detected": detected}
    
    async def extract_node(state: ScrapeState) -> dict:
        """LLM-driven extraction with strict JSON Schema."""
        schema = {
            "type": "object",
            "properties": {
                "title": {"type": "string"},
                "price": {"type": "number"},
                "currency": {"type": "string"},
                "in_stock": {"type": "boolean"},
            },
            "required": ["title", "price", "currency", "in_stock"],
            "additionalProperties": False,
        }
    
        resp = await client.chat.completions.create(
            model="gpt-4o-mini",
            response_format={
                "type": "json_schema",
                "json_schema": {"name": "product", "schema": schema, "strict": True},
            },
            messages=[
                {"role": "system", "content": "Extract product data from HTML."},
                {"role": "user", "content": (state["html"] or "")[:200000]},
            ],
        )
        return {"extracted": json.loads(resp.choices[0].message.content)}
    
    async def captcha_solver_node(state: ScrapeState) -> dict:
        """Stub — wire to 2Captcha, CapSolver, or similar."""
        return {
            "errors": state["errors"] + ["captcha solver not implemented"],
            "captcha_detected": False,
        }
    

    Wiring the graph

    This is the part that beats every alternative framework on clarity. You list nodes, list edges, and the runtime is built.

    def route_after_captcha_check(state: ScrapeState):
        if state["captcha_detected"]:
            return "captcha_solver"
        return "extract"
    
    def route_after_fetch(state: ScrapeState):
        if state.get("html") is None:
            if state["attempt"] >= 3:
                return END
            return "fetch"
        return "captcha_check"
    
    graph = StateGraph(ScrapeState)
    graph.add_node("fetch", fetch_node)
    graph.add_node("captcha_check", captcha_check_node)
    graph.add_node("captcha_solver", captcha_solver_node)
    graph.add_node("extract", extract_node)
    
    graph.set_entry_point("fetch")
    graph.add_conditional_edges("fetch", route_after_fetch)
    graph.add_conditional_edges("captcha_check", route_after_captcha_check)
    graph.add_edge("captcha_solver", "fetch")  # retry after solving
    graph.add_edge("extract", END)
    
    app = graph.compile()
    

    That graph handles the basic happy path, captcha branch, and a 3-attempt retry on fetch failures. LangGraph compiles it into an executable that you invoke with the initial state.

    async def main():
        final = await app.ainvoke({
            "url": "https://www.lazada.sg/products/xyz",
            "html": None,
            "captcha_detected": False,
            "extracted": None,
            "errors": [],
            "attempt": 0,
        })
        print(json.dumps(final["extracted"], indent=2))
    
    asyncio.run(main())
    

    Adding a checkpointer

    For long-running scraping jobs (think: scrape ten thousand products with intermediate state at each one), the checkpointer is mandatory. SQLite for development, Postgres for production.

    from langgraph.checkpoint.postgres import PostgresSaver
    
    DB_URI = "postgresql://scraper:secret@localhost:5432/scrapes"
    
    async with PostgresSaver.from_conn_string(DB_URI) as checkpointer:
        await checkpointer.setup()
        app = graph.compile(checkpointer=checkpointer)
    
        config = {"configurable": {"thread_id": "lazada-product-12345"}}
    
        async for event in app.astream(initial_state, config):
            print(event)
    

    If the worker dies mid-graph, you restart with the same thread_id and LangGraph resumes from the last successful node. Critical for any scrape that takes more than a minute.

    Adding tool nodes for proxy rotation

    For real production work, every fetch should go through a rotating proxy pool. Wire it into the fetch node:

    import random
    import os
    
    PROXIES = os.environ.get("PROXY_POOL", "").split(",")
    
    async def fetch_node(state: ScrapeState) -> dict:
        proxy = random.choice(PROXIES) if PROXIES and PROXIES != [""] else None
        proxy_config = None
        if proxy:
            u, _, rest = proxy.partition("://")
            if "@" in rest:
                creds, host_port = rest.split("@", 1)
                user, password = creds.split(":")
                proxy_config = {"server": f"{u}://{host_port}", "username": user, "password": password}
            else:
                proxy_config = {"server": proxy}
    
        try:
            async with async_playwright() as p:
                browser = await p.chromium.launch(headless=True, proxy=proxy_config)
                page = await browser.new_page()
                await page.goto(state["url"], wait_until="networkidle", timeout=30000)
                html = await page.content()
                await browser.close()
            return {"html": html, "attempt": state["attempt"] + 1}
        except Exception as e:
            return {
                "errors": state["errors"] + [f"fetch failed: {e}"],
                "attempt": state["attempt"] + 1,
            }
    

    For ASEAN scraping where mobile IPs work better than residential, Singapore mobile proxy integrates as the proxy pool source.

    Sticky proxies via thread state

    For multi-page flows where the same session must hold the same exit IP, store the proxy in the state and pin it across nodes:

    class ScrapeState(TypedDict):
        url: str
        html: Optional[str]
        captcha_detected: bool
        extracted: Optional[dict]
        errors: List[str]
        attempt: int
        sticky_proxy: Optional[str]  # bound on first fetch, reused across retries
    
    async def fetch_node(state: ScrapeState) -> dict:
        proxy = state.get("sticky_proxy") or random.choice(PROXIES)
        # ... fetch with proxy
        return {"html": html, "sticky_proxy": proxy, "attempt": state["attempt"] + 1}
    

    This is essential for cart and checkout flows on retailers that fingerprint the session-to-IP binding.

    Parallel fan-out with Send

    For the common pattern of “scrape 50 URLs and aggregate the results,” LangGraph supports a Send primitive that fans out across N parallel branches and rejoins.

    from langgraph.graph import Send
    
    def fanout(state):
        return [Send("scrape_one", {"url": u, "html": None, "errors": [], "attempt": 0})
                for u in state["urls"]]
    
    graph = StateGraph(BatchState)
    graph.add_node("scrape_one", scrape_one_node)
    graph.add_node("aggregate", aggregate_node)
    graph.set_entry_point("dispatcher")
    graph.add_conditional_edges("dispatcher", fanout, ["scrape_one"])
    graph.add_edge("scrape_one", "aggregate")
    graph.add_edge("aggregate", END)
    

    This gets you 50-way parallel scraping with proper backpressure (limit concurrency in the runtime config) and a single aggregated result. The pattern is faster than spawning 50 separate graph invocations because the aggregate state lives in one place.

    Retry strategies that actually work

    The naive retry pattern is “if fetch fails, retry up to N times.” In production this is rarely sufficient because failures cluster: a bad IP fails 5 times in a row before you decide to rotate. A better pattern uses exponential backoff and IP rotation between attempts.

    async def fetch_node(state: ScrapeState) -> dict:
        import asyncio
        backoff = min(2 ** state["attempt"], 30)
        if state["attempt"] > 0:
            await asyncio.sleep(backoff)
        proxy = random.choice(PROXIES)  # NEW proxy on every retry
        # ... fetch with proxy
    

    Combined with a circuit breaker on the proxy pool (evict any IP that fails 3 times in 10 minutes), the success rate on real-world targets jumps from roughly 88 percent to over 97 percent.

    Comparing LangGraph to alternatives

    Framework State model Branching Persistence Observability Best fit
    LangGraph Explicit TypedDict First-class Built-in checkpointer LangSmith integration Complex multi-step pipelines
    LangChain ReAct Implicit, chat history Prompt-driven Manual LangSmith integration Simple one-shot tasks
    CrewAI Per-crew shared memory Role-based Manual LangSmith or self Multi-agent role play
    AutoGen Group chat state Free-form Manual OpenTelemetry Conversational agents
    Custom asyncio Whatever you build Whatever you write Whatever you write Whatever you wire Maximum flexibility

    LangGraph wins for scraping specifically because scraping flows are state machines with clear branches: fetch then maybe captcha then extract then maybe retry. That maps to LangGraph’s primitives one to one. ReAct loops can do the same job but every behavior change requires prompt rewriting.

    For more on CrewAI as an alternative, see CrewAI for scraping pipelines. For AutoGen, see Multi-agent scraping with AutoGen in 2026.

    Cost benchmarks

    Single product page extraction, end to end, on the graph above:

    LLM model Avg LLM tokens per page LLM cost per page Compute per page Total per 1k pages
    GPT-4o-mini 14,000 $0.0028 $0.001 $3.80
    GPT-4o 14,000 $0.05 $0.001 $51
    Claude 3.5 Sonnet 13,500 $0.052 $0.001 $53
    Claude 3.5 Haiku 13,500 $0.011 $0.001 $12

    For high-volume scraping where you control the prompt and the schema is simple, GPT-4o-mini or Claude Haiku is the right pick. For tricky sites where extraction quality matters more than cost, Sonnet or 4o is worth it.

    Production deployment

    Run LangGraph workers under a process supervisor (systemd, PM2, or Kubernetes) with a health-check endpoint. Use Redis or Postgres for the checkpointer in production.

    A minimal worker loop:

    import asyncio
    from redis.asyncio import Redis
    
    redis = Redis.from_url("redis://localhost:6379")
    
    async def worker():
        while True:
            url = await redis.brpop("scrape:queue", timeout=10)
            if url is None:
                continue
            url = url[1].decode()
            state = make_initial_state(url)
            config = {"configurable": {"thread_id": f"scrape-{url}"}}
            try:
                final = await app.ainvoke(state, config)
                await store_result(url, final)
            except Exception as e:
                await redis.lpush("scrape:dead", url)
    
    asyncio.run(worker())
    

    For LangGraph deployment on serverless, the LangGraph Platform docs cover the managed option in detail.

    A complete production graph with all the layers

    Putting the patterns together gives a graph that handles the long tail. The nodes:

    1. validate_url rejects malformed input early, saving downstream work.
    2. fetch with sticky proxy and exponential backoff.
    3. captcha_check flags Cloudflare, DataDome, PerimeterX, and Akamai.
    4. captcha_solver calls 2Captcha for Turnstile and CapSolver for the rest.
    5. extract pulls structured fields with strict schema.
    6. validate_extracted checks Pydantic-level invariants (price > 0, in_stock is boolean).
    7. enrich adds derived fields (USD-converted price, normalized SKU).
    8. persist writes to Postgres with an upsert.
    9. notify posts to Slack on price changes greater than 10 percent.

    The graph branches at captcha_check (solver vs extract), at validate_extracted (re-fetch vs persist on validation failure), and at notify (skip if no significant change). Total node count: 9. Total edges: 14. The graph compiles in milliseconds and runs each scrape in roughly 4 to 8 seconds depending on captcha presence.

    Real teams underestimate how much of their scraper is the surrounding plumbing (validate, enrich, persist, notify) versus the core fetch and extract. Having those as named nodes in a graph instead of buried inside the fetch function makes the system far easier to reason about when something breaks at 2 AM.

    Cost and latency under realistic load

    Numbers from a March 2026 production deployment running 50,000 product scrapes per day on a 4 vCPU 8 GB worker pool:

    Metric Value
    Median graph latency end-to-end 4.8 s
    p99 graph latency 22 s
    Throughput per worker 12 scrapes/min
    Workers needed for 50k/day 3
    LLM cost per scrape (GPT-4o-mini) $0.0028
    Proxy cost per scrape (residential) $0.0006
    Compute cost per scrape (Fargate) $0.0008
    Total per-scrape cost $0.0042
    Daily total $210

    Compared to a non-LangGraph baseline (raw asyncio plus prompt) that ran at $190/day for the same throughput, LangGraph adds about 10 percent overhead in exchange for resumability, observability, and retry correctness. Most teams find the tradeoff lopsidedly worth it.

    Observability with LangSmith

    LangSmith remains the easiest way to see what a LangGraph agent is doing in production. Every node execution is traced with inputs, outputs, latency, and any LLM calls inside the node. The trace tree mirrors the graph structure, so you can spot a slow extract node or a node that errored without grepping logs.

    Three patterns worth adopting from the start:

    Tag every run with the URL being scraped and the source queue name. This makes filtering on the LangSmith UI trivial. langsmith_extra={"tags": [url_domain, queue_name]} in the invoke call.

    Add custom metadata for the scrape ID and the proxy used. When something fails on a specific IP, you can filter by proxy to see if the failure is sticky to one bad IP versus a target-side ban.

    Use LangSmith’s evaluation suite to regress against a frozen set of known-good HTML pages. Whenever you change the extract prompt, run the eval and confirm structured output quality has not regressed. This catches subtle prompt drift that production logs would not surface for days.

    For teams that prefer self-hosted observability, OpenTelemetry instrumentation is supported via the langsmith SDK with the OTel exporter. Span attributes match the LangGraph node names, so Jaeger or Tempo show the same trace tree.

    Frequently asked questions

    Can LangGraph handle parallel scraping of multiple URLs?
    Yes. Spawn one app invocation per URL, each with its own thread_id, and let asyncio handle concurrency. LangGraph also supports parallel branches inside a single graph using Send for fan-out patterns.

    How do I version a LangGraph workflow safely?
    Treat the graph definition as code, version it in git, and include a graph version in your thread_id. Old checkpoints stay tied to the old graph version, new ones to the new.

    Does LangGraph work with local LLMs?
    Yes. Anything that exposes an OpenAI-compatible API (Ollama, vLLM, LM Studio) plugs in as the LLM. Quality depends on the model. Llama 3.3 70B and Qwen 2.5 72B are the strongest open-source picks for extraction tasks in 2026.

    Can I use LangGraph from JavaScript?
    Yes. LangGraph.js is feature-equivalent to the Python version as of late 2025 and is the right pick for Node.js scraping pipelines.

    What about cycles? Can a graph loop forever?
    Cycles are allowed and useful for retry patterns, but every graph has a recursion_limit config (default 25) that prevents infinite loops. Set it explicitly for your use case.

    Can I share state between two unrelated graph invocations?
    Yes, by writing to an external store (Redis, Postgres) from inside a node. LangGraph state is per-thread by design; cross-thread sharing is intentional code, not implicit behavior.

    How do I add human-in-the-loop approval to a scraping graph?
    Use the interrupt_before config to pause the graph before a sensitive node (say, the one that publishes to your warehouse). LangGraph blocks until you call update_state and resume. This is how teams add manual review gates without restructuring the graph.

    What about streaming partial results to a client?
    The astream_events API emits a stream of node-level events. For long-running scrapes that feed a UI, stream the events over a WebSocket and the UI shows progress as each node completes.

    Does LangGraph have built-in support for batching multiple URLs into one LLM call?
    No, but you can implement it as a node that buffers up to N URLs and emits a single multi-extraction request. The trade-off is latency for cost: batching cuts LLM tokens by 30 to 50 percent on small extractions but adds a buffering delay. Most teams skip it and accept the per-page cost.

    Production gotchas

    • The Postgres checkpointer needs an explicit setup() call on first use. Skipping it produces a confusing “relation does not exist” error.
    • Conditional edges that return a list of node names cause parallel execution. Returning a single string is sequential routing. Mixing the two is the most common bug we see in code review.
    • LangGraph state is merged with a shallow update. If a node returns {"errors": [new_error]}, it overwrites the prior errors list. Use a reducer to append: errors: Annotated[list, operator.add].
    • recursion_limit counts node executions, not loop iterations. A complex graph with many parallel branches can hit the default 25 limit unexpectedly.
    • The Sqlite checkpointer is fine for development but locks aggressively under concurrent writes. Switch to Postgres before going production.
    • Returning a partial state from a node with no fields removed but without including unchanged fields is correct, and merging is automatic. New developers often re-emit the full state thinking they have to, which works but obscures intent.
    • The compile step is cheap; recompile on every code change in dev. In prod, compile once at startup and reuse the compiled app across requests.

    If you are building a scraping team in 2026, the AI agentic proxies category covers the proxy and infrastructure side that pairs with LangGraph for full production deployments.