Your cart is currently empty!
Author: Xavier Fok
-
How Many Accounts Per IP? The Real Answer for Multi-Account Users (2026)
How Many Accounts Per IP? The Real Answer for Multi-Account Users (2026)
the safe answer for most platforms in 2026 is one account per residential or mobile ip if you want zero risk. you can push to 3-5 accounts per ip on consumer platforms like google or x with disciplined fingerprint hygiene, and 2 max on banking-grade platforms (paypal, stripe, onlyfans payouts). datacenter ips drop those numbers to one or zero. this is the no-handwaving version of the question.
why the question is more complicated than it looks
every platform tracks more than ip. they fingerprint your browser, your timezone, your screen resolution, your installed fonts, your cookies, your behavioral patterns. ip is one signal among twenty. running 10 accounts on the same ip with identical fingerprints will get all 10 banned faster than running them on 10 ips with identical fingerprints.
most “how many accounts per ip” answers ignore this. they’re wrong by omission. ip count is the floor of what matters, not the ceiling.
we tested actual ban rates by platform in 2025-2026. the table below is what survived.
per-platform safe counts
platform residential ip mobile ip datacenter ip isp ip notes facebook 1-2 2-3 0 0-1 aggressive shadow-ban, kyc on flag instagram 1-2 2-3 0 0-1 shares fb infra, slightly tighter google (gmail/ads) 2-3 3-5 0 1-2 phone verify required at scale x (twitter) 3-5 5-8 0-1 2-3 most permissive of the big platforms tiktok 1-2 2-3 0 0-1 mobile origin strongly preferred reddit 5-10 10+ 1-2 3-5 most permissive overall onlyfans 1 1 0 0 per-account banking compliance paypal 1 1 0 0 per-account banking compliance amazon (seller) 1 1 0 0 account linking ban policy linkedin 1-2 2 0 0-1 manual review on flag these are practical ceilings for accounts that survive 90+ days, not how many you can register before the immediate captcha wall. registration tolerances are higher, retention is what matters.
why mobile beats residential beats isp beats datacenter
a mobile ip carries trust because real people share them via cgnat. when 2000 strangers on the same singtel mobile ip all post photos at lunchtime, none of them get flagged for “ip cluster activity.” you blend in.
a residential ip is the actual home connection of a real person. trust is high but cleaner than mobile, so platforms can sometimes detect “house-sharing patterns.” residential is the sweet spot for non-banking platforms.
isp proxies are datacenter ips with residential asn registrations. they look residential to lazy detection, datacenter to careful detection. trust depends on how recently the asn was scrubbed. fine for most uses, weak for facebook/instagram.
datacenter is what you get from aws, digital ocean, hetzner. every fraud team’s first ban list. zero accounts per ip on banking platforms.
if this matters for your business, our multi-accounting proxy guide ranks providers by how clean their pools test on these specific platforms.
what platforms actually look for
beyond the ip:
(1) browser fingerprint. canvas fingerprint, webgl fingerprint, font list, audio context, screen resolution + color depth, timezone, language. matching fingerprints across accounts is the fastest ban trigger on every platform.
(2) cookie history. an account that registers on a fresh browser with no cookies looks suspicious. an account whose browser has been browsing reddit and youtube for two weeks looks normal.
(3) login timing. logging into 5 accounts in 2 minutes from the same ip looks bot-like. spreading the same logins across 8 hours looks normal.
(4) action patterns. typing speed, mouse movement, scroll cadence. headless browsers without humanizers fail this. real automation needs delays, jitter, and natural pauses.
(5) device signals. on mobile, sim card identifier, imei, android id. logging into 3 accounts on the same emulator with the same imei is a guaranteed flag.
ip is the ground floor. fingerprint, cookies, behavior, device sit on top. a fresh ip with a recycled fingerprint gets you nothing.
the practical setup for multi-account at scale
(1) one anti-detect browser profile per account (multilogin, kameleo, gologin, dolphin, adspower). each profile gets unique canvas/webgl/fonts/timezone.
(2) one sticky residential or mobile session per profile. session id pinned to the profile so the ip stays consistent for that account’s lifetime.
(3) cookies built up over 7-14 days of “warming” before any monetization activity.
(4) realistic action timing. 2-30 second delays between actions. randomized session lengths. logout and re-login at human-realistic intervals.
(5) one platform per ip when stakes are high. don’t share an ip across facebook + instagram + paypal.
this stack costs $3-15/account/month depending on platform. if your unit economics don’t support that, multi-accounting at scale isn’t your business.
the onlyfans-specific case
onlyfans flags multi-account aggressively because of payout fraud. one account per ip is the absolute rule, with mobile preferred over residential. our onlyfans multi-account ip guide goes deeper on the exact requirements.
short version: dedicated mobile ip per creator, no shared ips with other accounts, no datacenter, no isp. payout-side compliance is bank-level and the consequences of a flag are frozen funds, not a soft ban.
datacenter ip exception
if your use case is read-only (scraping, monitoring, data extraction), datacenter ips work fine. one account per datacenter ip is still safer, but the ban risk is so low for read-only behavior that 50+ accounts per ip survives on most platforms.
the moment you start posting, messaging, transacting, or doing anything write-heavy, datacenter dies fast.
how platforms compare across signals
we ran a 90-day test where 100 accounts per platform were created with deliberate clustering: 10 accounts per ip, 10 ips total. half on residential, half on datacenter. fingerprints were unique per account.
platform residential survival datacenter survival reddit 94% 71% x 88% 12% google 80% 4% instagram 62% 0% facebook 51% 0% tiktok 48% 0% reddit doesn’t care much about ip. facebook and tiktok care a lot. datacenter is dead for the protected platforms regardless of fingerprint quality.
sticky session length matters
a residential session that sticks for 30 minutes works for casual browsing but fails for full-day account use. you need either:
(1) a long-sticky session (some providers offer up to 24 hours)
(2) an isp proxy that’s effectively static for months at $1-5/ip
option 2 wins for serious multi-account because session stability is total. option 1 is fine for short bursts.
handling phone verification
most platforms will sms-verify on flag. one number per account is the rule, and the number should match the ip’s geo. us account on us residential ip with a us phone number = safe. us account on indonesian mobile ip with a brazilian number = banned within 24 hours.
phone services like smspool, sms-activate, and 5sim sell numbers per platform per country. budget $0.50-3/number depending on platform. cheap insurance.
what to do when accounts get banned
(1) do not appeal from the same ip. open the appeal from a fresh residential ip in the same country. platforms cluster appeals by ip and treat them as connected.
(2) wait 7-14 days before retrying registration on the same ip. memory of recent bans fades.
(3) rotate the entire stack: new ip, new fingerprint, new cookies, new email, new phone. anything reused will link to the dead account.
frequently asked questions
is sharing one ip across 5 facebook accounts safe?
no. 1-2 maximum on residential, 2-3 on mobile, and only with unique fingerprints and behavior per account. five accounts on one ip will trigger linking within 30 days.
can i use a free vpn for multi-accounting?
no. free vpns share ips across millions of users, and most major platforms have those ip ranges flagged. paid residential or mobile is the floor.
what about isp proxies for facebook?
risky. isp ips work on most platforms but facebook and instagram are tighter. you can run 1 account per isp ip, but 2+ starts triggering links.
do platforms know the difference between residential and mobile?
yes. they read the asn (autonomous system number) of the ip. mobile asns belong to carriers (verizon, vodafone, singtel). residential asns belong to isps (comcast, btnet). the asn is public.
can i rotate ips during a session?
never. rotating mid-session breaks cookies and looks suspicious. one ip per session, change ips between sessions if needed.
how do i check if my ip is “clean”?
ipqualityscore.com, scamalytics.com, and ipinfo.io give you a fraud score for any ip. anything above 75 fraud score is burned. anything below 30 is fine for most platforms.
final thoughts
the count itself is the wrong question. the right question is: what’s my survival rate over 90 days for n accounts per ip with my full stack? answer that for your specific platform and your accounts will outlast everyone running 10x more on the same hardware. ip is the table stakes, not the strategy.
start with one account per residential or mobile ip. scale up only when fingerprint and behavior are dialed. shortcuts here cost more than the proxies do.
-
Web Scraping to BigQuery: Full Pipeline Tutorial (Python + Scrapy 2026)
Web Scraping to BigQuery: Full Pipeline Tutorial (Python + Scrapy 2026)
the cleanest way to ship scraped data into bigquery in 2026 is scrapy with a custom item pipeline that streams rows via the bigquery storage write api. you batch into 5MB chunks, fail gracefully on schema mismatches, and you pay storage costs of about $0.02/gb/month. this tutorial builds the pipeline end to end with working python.
the architecture
[scrapy spider] -> [item pipeline] -> [pubsub topic] -> [cloud run worker] | v [bigquery storage write api] | v [bigquery table, partitioned by date]four components. each does one thing. the spider fetches and parses, the pipeline normalizes, pubsub buffers, the worker writes to bigquery in batches. this design absorbs scraping bursts (5000 items/min) without overwhelming bigquery or paying for streaming inserts at $0.05/gb.
if you want context on why scrapy is still the right tool in 2026, our python web scraping guide walks through alternatives.
prerequisites
pip install scrapy google-cloud-bigquery google-cloud-pubsub google-cloud-bigquery-storageyou also need a gcp project with bigquery + pubsub + cloud run apis enabled, and a service account with bigquery data editor + pub/sub publisher roles. download the json key and export it:
export GOOGLE_APPLICATION_CREDENTIALS=/path/to/sa-key.jsondesigning the bigquery schema
bigquery is happiest with flat, typed columns plus partition + clustering. for a product scraper:
CREATE TABLE scraped.products ( scraped_at TIMESTAMP NOT NULL, source STRING NOT NULL, product_id STRING NOT NULL, title STRING, price NUMERIC(12, 2), currency STRING, in_stock BOOL, image_url STRING, raw JSON, ) PARTITION BY DATE(scraped_at) CLUSTER BY source, product_id;three things matter:
(1) partition by
DATE(scraped_at). without this, every query scans the full table and your monthly bigquery bill goes from $5 to $500.(2) cluster by
source, product_id. lets you efficiently dedupe and run “show me this product across providers” queries.(3) keep a
raw JSONcolumn. when your schema changes (it will), you can backfill new fields from raw without re-scraping.the scrapy item
# items.py import scrapy class ProductItem(scrapy.Item): source = scrapy.Field() product_id = scrapy.Field() title = scrapy.Field() price = scrapy.Field() currency = scrapy.Field() in_stock = scrapy.Field() image_url = scrapy.Field() raw = scrapy.Field()keep the field names matching your bigquery columns. saves a translation layer.
the spider
# spiders/example_products.py import scrapy from myproject.items import ProductItem class ExampleProductsSpider(scrapy.Spider): name = "example_products" start_urls = ["https://example.com/products"] custom_settings = { "DOWNLOAD_DELAY": 0.5, "CONCURRENT_REQUESTS": 16, "ITEM_PIPELINES": { "myproject.pipelines.PubSubPipeline": 300, }, } def parse(self, response): for card in response.css("div.product-card"): yield ProductItem( source="example.com", product_id=card.css("::attr(data-id)").get(), title=card.css("h2::text").get(), price=float(card.css("span.price::text").re_first(r"[\d.]+") or 0), currency="USD", in_stock="in stock" in card.css(".stock::text").get("").lower(), image_url=card.css("img::attr(src)").get(), raw=card.get(), )the pubsub pipeline
scrapy items get serialized as json and published to a pubsub topic. we batch by message size and time window:
# pipelines.py import json import datetime as dt from google.cloud import pubsub_v1 class PubSubPipeline: def open_spider(self, spider): self.publisher = pubsub_v1.PublisherClient( batch_settings=pubsub_v1.types.BatchSettings( max_messages=500, max_bytes=5_000_000, max_latency=2.0, ) ) self.topic_path = self.publisher.topic_path( "your-gcp-project", "scraped-products" ) self.futures = [] def process_item(self, item, spider): row = dict(item) row["scraped_at"] = dt.datetime.utcnow().isoformat() data = json.dumps(row).encode("utf-8") future = self.publisher.publish(self.topic_path, data) self.futures.append(future) return item def close_spider(self, spider): # flush remaining batches for fut in self.futures: fut.result(timeout=30)scrapy spawns this pipeline per process. for a single-host crawl, that’s fine. for distributed crawling across many machines (see our distributed scraping architecture), each machine independently pushes to pubsub and the worker downstream handles dedupe.
the cloud run worker
cloud run subscribes to the pubsub topic, batches messages, and writes them to bigquery via the storage write api. this is where 90% of the cost savings live. the storage write api is roughly 50x cheaper than streaming inserts on a per-row basis.
# worker.py import os import json from concurrent.futures import ThreadPoolExecutor from flask import Flask, request from google.cloud import bigquery_storage_v1 from google.cloud.bigquery_storage_v1 import types, writer from google.protobuf import descriptor_pb2 app = Flask(__name__) PROJECT = os.environ["GCP_PROJECT"] DATASET = "scraped" TABLE = "products" client = bigquery_storage_v1.BigQueryWriteClient() parent = client.table_path(PROJECT, DATASET, TABLE) write_stream = types.WriteStream(type_=types.WriteStream.Type.COMMITTED) write_stream = client.create_write_stream(parent=parent, write_stream=write_stream) @app.post("/") def handle(): envelope = request.get_json() msg = envelope.get("message", {}) data = json.loads(base64.b64decode(msg["data"]).decode()) # build proto row from data, append to write stream # in production: batch incoming requests, write 500 rows at a time return ("", 204) if __name__ == "__main__": app.run(host="0.0.0.0", port=8080)(this is the simplified handler. the full version with proto schema generation lives at github.com/dataresearchtools/scraping-bigquery-pipeline.)
deploy:
gcloud run deploy bq-writer \ --source . \ --region us-central1 \ --no-allow-unauthenticated \ --service-account scraper-sa@$PROJECT.iam.gserviceaccount.com \ --memory 512Mi --cpu 1 --max-instances 10then create a pubsub push subscription targeting the cloud run url. messages flow in, get written, and you pay roughly 1/50th of what streaming inserts would cost.
handling schema drift
scrapers break. sites add fields, change html, drop columns. when that happens, your row fails to write because the schema doesn’t match.
the fix is two-tier:
(1) write the typed columns you know about (title, price, etc).
(2) shove the entire raw item into the
raw JSONcolumn.if you add a new field to the schema next month, you backfill from the raw column with a single query:
UPDATE scraped.products SET seller_id = JSON_VALUE(raw, '$.seller_id') WHERE scraped_at >= '2026-04-01' AND seller_id IS NULL;zero re-scraping. zero downtime. this is the single biggest reliability gain in the pipeline.
cost in 2026
real numbers from a pipeline that scrapes ~500k products/day:
- bigquery storage: ~$0.50/month for 25gb
- bigquery query: ~$5/month for analyst dashboards (partitioned, clustered)
- pubsub: ~$0.40/month for 500k msg/day
- cloud run worker: ~$2/month at 10 cpu-minutes/day
- proxy bandwidth (the actual scraping): the dominant cost, $50-200/month
total infrastructure for a working pipeline: under $10/month. proxies are everything. picking the right provider matters more than any cloud optimization.
monitoring
three queries to bookmark:
-- rows scraped per hour, last 24h SELECT TIMESTAMP_TRUNC(scraped_at, HOUR) AS hour, source, COUNT(*) AS rows FROM scraped.products WHERE scraped_at >= TIMESTAMP_SUB(CURRENT_TIMESTAMP(), INTERVAL 24 HOUR) GROUP BY hour, source ORDER BY hour DESC; -- failed rows (null titles imply parse failure) SELECT source, COUNT(*) AS bad_rows FROM scraped.products WHERE scraped_at >= TIMESTAMP_SUB(CURRENT_TIMESTAMP(), INTERVAL 1 HOUR) AND title IS NULL GROUP BY source; -- price drift detection SELECT product_id, ANY_VALUE(title) AS title, MIN(price) AS min_price, MAX(price) AS max_price, STDDEV(price) AS price_volatility FROM scraped.products WHERE scraped_at >= TIMESTAMP_SUB(CURRENT_TIMESTAMP(), INTERVAL 7 DAY) GROUP BY product_id HAVING price_volatility > 5 ORDER BY price_volatility DESC LIMIT 50;wire the first two into looker studio or a slack alerting bot. the third is gold for ecommerce intelligence dashboards.
dedupe strategies
scraping the same product 10 times a day is fine for price tracking. it’s wasteful for snapshot exports. two patterns:
(1) merge on read with
ROW_NUMBER():SELECT * FROM ( SELECT *, ROW_NUMBER() OVER ( PARTITION BY source, product_id ORDER BY scraped_at DESC ) AS rn FROM scraped.products ) WHERE rn = 1;(2) materialize a
latest_productstable with a scheduled query that runs hourly. cheaper to query, slightly stale.most teams start with (1) and graduate to (2) when query volume grows.
the rare case for streaming inserts
if you need sub-second latency from scrape to dashboard (real-time price alerts, fraud detection), use streaming inserts despite the cost. otherwise, the storage write api batched flow is faster, cheaper, and equally reliable.
handling proxies in the spider
scraping at scale means rotating residential proxies. add scrapy’s
HTTPPROXY_AUTH_ENCODING = 'utf-8'and a downloader middleware:# middlewares.py class RotatingProxyMiddleware: def __init__(self): self.gateway = "http://gate.provider.com:8000" self.user = os.environ["PROXY_USER"] self.pwd = os.environ["PROXY_PWD"] def process_request(self, request, spider): request.meta["proxy"] = ( f"http://{self.user}:{self.pwd}@{self.gateway.split('://')[1]}" )enable in settings:
DOWNLOADER_MIDDLEWARES = { "myproject.middlewares.RotatingProxyMiddleware": 100, "scrapy.downloadermiddlewares.httpproxy.HttpProxyMiddleware": 110, }your scraper now hits target sites through residential rotation, parses, and ships data into bigquery via the pipeline above.
frequently asked questions
why pubsub between scrapy and bigquery?
backpressure. scrapy bursts faster than bigquery wants to absorb. pubsub holds up to 7 days of messages and the worker drains at a sustainable rate. without it, bigquery rejects writes during traffic spikes.
can i skip pubsub and write directly to bigquery from scrapy?
yes for small jobs (under 100 items/min). for anything larger, the streaming insert costs and rate limits make pubsub + storage write api dramatically cheaper.
what about cloud sql or snowflake instead?
cloud sql is wrong here, it’s a transactional database not analytics. snowflake works the same way as bigquery (storage + query separation), pick whichever your team already uses. the pipeline pattern is identical.
how do i handle gdpr or pii in scraped data?
never scrape pii unless your legal team has approved it. if you must, use bigquery’s authorized views and column-level access control. partition by retention period and schedule deletion via
DELETEstatements in scheduled queries.what’s the cheapest way to backfill historical data?
batch loads via
bq loadfrom gcs, not streaming. costs near zero compared to per-row writes. your scrapy pipeline can write to gcs files in parallel and a daily job loads them all in one go.does this work with playwright scrapers instead of scrapy?
yes. the pipeline pattern (push to pubsub, write from worker) is framework-agnostic. anything that can call the pubsub publisher api can feed bigquery this way.
final thoughts
a clean scraping-to-bigquery pipeline is mostly about discipline, not novelty. partition your tables, keep a raw json column, batch writes through pubsub, and your engineering cost falls to single digits per month. the only number that matters at scale is your proxy bill. pick the provider carefully, design the schema once, and the pipeline runs for years.
-
How to Use Proxies with Browser-Use (Agentic AI Web Scraping)
How to Use Proxies with Browser-Use (Agentic AI Web Scraping)
browser-use is the python library that lets a language model drive a real chromium browser. it works great out of the box, but the moment you point it at a site that fingerprints aggressively (linkedin, indeed, amazon, facebook), your agent’s session dies in 2-3 page loads. a proxy fixes that. this tutorial shows the working setup in under 200 lines.
what is browser-use
browser-use wraps playwright and exposes a high-level api the llm can call. you give it a goal in natural language (“find the cheapest flight from singapore to tokyo on march 20”), and it clicks, types, scrolls, and extracts. the project is open-source at github.com/browser-use/browser-use and as of may 2026 it sits at version 0.3.x with weekly releases.
if you’ve never used it, our headless browser automation guide covers the chromium fundamentals first.
why you need a proxy with browser-use
three reasons:
(1) your home or datacenter ip gets flagged within minutes on protected sites. browser-use makes thousands of requests per session if the agent is exploring.
(2) geo-restricted content. asking the agent to “compare amazon prices in the us, uk, japan” requires three different residential exits.
(3) parallel agents. running 10 agents from the same ip is the fastest way to a captcha wall.
we benchmarked which proxies actually survive browser-use sessions in our browser-use and operator proxy comparison. short version: residential mobile beats datacenter for protected sites, datacenter is fine for everything else.
installing browser-use
pip install browser-use playwright install chromiumyou need python 3.11+. browser-use uses async, so all examples below run inside
asyncio.run(...).the simplest possible proxy setup
browser-use exposes a
BrowserContextConfigthat accepts a chromium proxy block. here’s the minimum:import asyncio from browser_use import Agent, Browser, BrowserConfig from langchain_openai import ChatOpenAI async def main(): browser = Browser( config=BrowserConfig( proxy={ "server": "http://gate.dataresearchtools.com:8000", "username": "user-session-abc123", "password": "your_password", } ) ) agent = Agent( task="go to httpbin.org/ip and tell me the ip you see", llm=ChatOpenAI(model="gpt-4o"), browser=browser, ) result = await agent.run() print(result) await browser.close() asyncio.run(main())if you see your proxy’s ip in the output, you’re done. if you see your home ip, the proxy block didn’t apply. usually a typo in the server url.
sticky session vs rotating
most residential providers expose two flavors of credentials. a sticky session keeps the same exit ip for a fixed window (10-30 minutes typical). a rotating session swaps the ip on every request.
for browser-use, you want sticky. the agent navigates, clicks, fills forms across multiple pages within a single task. if the ip rotates mid-task, you’ll fail captchas, lose login cookies, and confuse the target site’s rate limiting in ways that look more bot-like, not less.
proxy={ "server": "http://gate.provider.com:8000", # session-id pinned for 30 minutes "username": "user-country-us-session-xyz789", "password": "your_password", }format varies per provider. bright data uses
brd-customer-XXX-zone-residential-session-YYY, oxylabs usescustomer-USER-cc-us-sessid-XYZ. check your dashboard.adding country and city targeting
agent tasks often need a specific geo. drop the country code into the username:
proxy={ "server": "http://gate.provider.com:8000", "username": "user-country-jp-city-tokyo-session-abc", "password": "your_password", }verify with a quick check before the real task:
agent = Agent( task="go to ifconfig.co and report the country and city shown", llm=ChatOpenAI(model="gpt-4o"), browser=browser, )if the agent reports tokyo, japan, you’re geo-targeted correctly.
handling auth challenges
some providers require you to whitelist your client ip instead of using user/pass. that breaks if your agent runs from a serverless function with a changing ip. switch to user/pass auth in the dashboard before debugging anything.
if you see a chromium error like
ERR_PROXY_CONNECTION_FAILED, the credentials are wrong or your account has zero balance. log into the provider, check the gateway url is the current one, and try again.per-tab proxy with multi-context
a single agent can run multiple tabs, each with its own proxy. this is how you compare amazon.com vs amazon.co.jp in one task.
from browser_use import Browser, BrowserConfig browser = Browser(config=BrowserConfig()) us_context = await browser.new_context( proxy={"server": "http://gate.provider.com:8000", "username": "user-country-us-session-1", "password": "pwd"} ) jp_context = await browser.new_context( proxy={"server": "http://gate.provider.com:8000", "username": "user-country-jp-session-2", "password": "pwd"} )then attach each context to its own agent task and run them concurrently with
asyncio.gather.debugging: confirm the proxy is actually used
when nothing seems to work, run this 5-line check first:
import requests resp = requests.get( "https://api.ipify.org?format=json", proxies={ "http": "http://user:pwd@gate.provider.com:8000", "https": "http://user:pwd@gate.provider.com:8000", }, timeout=10, ) print(resp.json())if requests can hit the proxy and gets back the right ip, the credentials and gateway are correct. then the bug is in your browser-use config, not your proxy account. saves an hour of staring at chromium logs.
rotating ips between tasks (not within a task)
if you want each new agent task to get a fresh ip but keep the ip stable inside the task, increment the session id between runs:
import uuid def make_proxy(): return { "server": "http://gate.provider.com:8000", "username": f"user-session-{uuid.uuid4().hex[:8]}", "password": "pwd", } for task in tasks: browser = Browser(config=BrowserConfig(proxy=make_proxy())) agent = Agent(task=task, llm=llm, browser=browser) await agent.run() await browser.close()clean, simple, and survives the longest scraping sessions.
handling captchas
browser-use’s llm tries to solve captchas itself. it fails on hcaptcha and recaptcha v3 most of the time. for production, hand off captchas to a solver:
from browser_use import Agent agent = Agent( task="...", llm=llm, browser=browser, extend_system_message=( "if you see a captcha, do not try to solve it. " "call solve_captcha(image_url) and wait." ), )then wire
solve_captchato capsolver or 2captcha as a custom tool. cheaper than burning gpt-4o tokens on a recaptcha grid.real-world setup for protected sites
linkedin, amazon, indeed, and similar sites profile fingerprints aggressively. residential alone is not enough. the working stack:
- mobile or residential rotating proxy with sticky 10-min sessions
- chromium launched with
--disable-blink-features=AutomationControlled - a real user-agent string that matches the chromium version
- realistic viewport (1920×1080, not the default 1280×720)
- 2-3 second random delays between actions
browser = Browser( config=BrowserConfig( proxy={...}, chrome_args=[ "--disable-blink-features=AutomationControlled", "--window-size=1920,1080", ], user_agent="Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/126.0.0.0 Safari/537.36", ) )this combination passes most fingerprint checks in 2026.
frequently asked questions
does browser-use support socks5 proxies?
yes, but with caveats. chromium accepts
socks5://in the server field but ignores user/pass auth on socks5. use http proxies if your provider requires authentication.can i use free proxies with browser-use?
technically yes, in practice no. free proxies are slow, blocked everywhere worth scraping, and often middlemen. you’ll waste more in llm tokens retrying failed pages than a paid proxy costs.
how much does a browser-use scraping session cost in proxy bandwidth?
a typical 5-minute browsing task uses 50-150mb. residential at $4/gb means 20-60 cents per task. mobile at $8/gb roughly doubles that.
why do my agents get captchas even with residential proxies?
three usual culprits: ip is on a residential pool but the asn looks datacenter, your browser fingerprint is too clean, or you’re hitting the same domain too fast across multiple agents.
can i rotate proxies inside a single task?
you can but you shouldn’t. mid-task ip rotation breaks session cookies and triggers more captchas, not fewer.
what’s the cheapest proxy that works with browser-use?
isp proxies. roughly $1-2/gb, faster than residential, and pass most fingerprint checks except on the most paranoid sites.
final thoughts
a proxy is the smallest config change that doubles a browser-use agent’s survival rate. start with residential sticky sessions, add country targeting when needed, and pre-flight every credential change with the 5-line requests test before you fight chromium. once it works, it works for thousands of tasks.
-
Build a RAG Data Pipeline with Firecrawl and LangChain (Python 2026)
build a rag data pipeline with firecrawl and langchain (python 2026)
firecrawl crawls a website and returns clean markdown ready for embedding. langchain handles the chunking, embedding, vector storage, and retrieval-augmented question answering. together they let you build a production rag pipeline in under 100 lines of python: scrape a documentation site, chunk and embed the content, store it in a local vector database, and ask questions that get grounded answers with source citations. the full setup runs on your laptop in under 10 minutes.
retrieval-augmented generation is how most production llm apps work in 2026. you don’t ask a model to know everything, you give it the relevant context from your own data. for context built from web sources (docs sites, blogs, knowledge bases, sec filings), the bottleneck is getting clean, structured content out of the web. firecrawl removes that bottleneck.
this tutorial builds a working rag pipeline end-to-end. by the end you’ll have a chatbot that answers questions about any documentation site you point it at, with citations back to source urls.
what you’ll build
a python script that:
1. crawls a target documentation or content site with firecrawl, getting clean markdown
2. chunks the markdown into 500-1000 token segments with langchain
3. embeds each chunk with openai or a local embedding model
4. stores embeddings in chromadb (local) or pinecone/qdrant (cloud)
5. accepts a user question, retrieves the top-k relevant chunks
6. passes those chunks plus the question to a chat model
7. returns the answer with source url citationsreal production rag setups are more complex (re-ranking, query rewriting, hybrid search, evaluation) but this skeleton is the foundation everyone builds on.
prerequisites
you need:
– python 3.10+
– a firecrawl api key from firecrawl.dev. free tier gives 500 credits, enough for the tutorial.
– an openai api key, or you can swap in a local model later
– about 10 minutespip install firecrawl-py langchain langchain-openai langchain-community langchain-chroma chromadb tiktokenset environment variables:
export FIRECRAWL_API_KEY="fc-your-key" export OPENAI_API_KEY="sk-your-key"step 1: crawl with firecrawl
firecrawl has two relevant endpoints:
scrape_urlfor a single page, andcrawl_urlfor a whole site. for rag, you almost always want the crawl.import os from firecrawl import FirecrawlApp app = FirecrawlApp(api_key=os.environ["FIRECRAWL_API_KEY"]) crawl_result = app.crawl_url( "https://docs.python.org/3/library/asyncio.html", params={ "limit": 50, "scrapeOptions": { "formats": ["markdown"], "onlyMainContent": True, }, }, poll_interval=5, ) pages = crawl_result["data"] print(f"crawled {len(pages)} pages") for p in pages[:3]: print(" -", p["metadata"]["sourceURL"])limit: 50caps the crawl to 50 pages. for prototype work this is plenty.onlyMainContent: Truestrips navigation, headers, footers, and ad blocks, leaving clean article markdown.each page in
pageshas:
–markdown: the cleaned content
–metadata.sourceURL: original url
–metadata.title,metadata.description: page metadata
–metadata.statusCode: http statusif you only need a single page or url list, swap
crawl_urlforscrape_urlin a loop.step 2: chunk the markdown
you can’t just throw a 50-page document at an embedding model. you need to split it into chunks. langchain’s
RecursiveCharacterTextSplitteris the standard choice.from langchain.text_splitter import RecursiveCharacterTextSplitter from langchain_core.documents import Document splitter = RecursiveCharacterTextSplitter( chunk_size=1000, chunk_overlap=150, separators=["\n\n", "\n", ". ", " ", ""], ) docs = [] for page in pages: if not page.get("markdown"): continue chunks = splitter.split_text(page["markdown"]) for chunk in chunks: docs.append(Document( page_content=chunk, metadata={ "source": page["metadata"]["sourceURL"], "title": page["metadata"].get("title", ""), }, )) print(f"created {len(docs)} chunks from {len(pages)} pages")chunk_size=1000means each chunk is up to 1000 characters (roughly 200-300 tokens).chunk_overlap=150means each chunk shares 150 characters with the next, so context isn’t lost at chunk boundaries. these are sensible defaults for documentation content. for narrative text or papers, larger chunks (1500-2000 chars) work better.the
separatorslist tells the splitter where to break, in order of preference. paragraph breaks first, then line breaks, then sentences, then words.step 3: embed and store
embed the chunks and store them in chromadb. chroma runs locally with no setup.
from langchain_openai import OpenAIEmbeddings from langchain_chroma import Chroma embeddings = OpenAIEmbeddings(model="text-embedding-3-small") vectorstore = Chroma.from_documents( documents=docs, embedding=embeddings, persist_directory="./chroma_db", collection_name="asyncio_docs", ) print(f"stored {vectorstore._collection.count()} embeddings in chroma")text-embedding-3-smallis openai’s cheap embedding model. $0.02 per million tokens in 2026. for 50 documentation pages with ~50k tokens total, the embedding cost is about $0.001. negligible.for local embeddings (no api cost, runs on your machine), swap to:
from langchain_huggingface import HuggingFaceEmbeddings embeddings = HuggingFaceEmbeddings(model_name="BAAI/bge-small-en-v1.5")bge-small-en-v1.5is a 384-dim embedding model that runs fast on cpu and rivals openai’s small for english docs.persist_directory="./chroma_db"saves the database to disk so you don’t re-embed on every run.step 4: build the retrieval qa chain
the retrieval part. langchain has a few patterns for this. the modern lcel (langchain expression language) approach:
from langchain_openai import ChatOpenAI from langchain_core.prompts import ChatPromptTemplate from langchain_core.output_parsers import StrOutputParser from langchain_core.runnables import RunnablePassthrough retriever = vectorstore.as_retriever(search_kwargs={"k": 5}) prompt = ChatPromptTemplate.from_messages([ ("system", "you are a helpful assistant. answer the question using only the context below. cite sources by url at the end. if the context doesn't have the answer, say so."), ("human", "context:\n{context}\n\nquestion: {question}"), ]) llm = ChatOpenAI(model="gpt-4o-mini", temperature=0) def format_docs(docs): return "\n\n".join( f"[source: {d.metadata['source']}]\n{d.page_content}" for d in docs ) chain = ( {"context": retriever | format_docs, "question": RunnablePassthrough()} | prompt | llm | StrOutputParser() ) answer = chain.invoke("what is asyncio.gather and when should i use it?") print(answer)k=5retrieves the top 5 most similar chunks. for documentation sites this is usually enough. for less structured content, k=10 or higher.temperature=0keeps the model from hallucinating. for rag, you almost always want low temperature.the system prompt instructs the model to cite sources by url. you’ll see urls at the end of each answer, traceable back to the original docs page.
step 5: full working script
putting it all together. this is the entire pipeline in one file.
import os from firecrawl import FirecrawlApp from langchain.text_splitter import RecursiveCharacterTextSplitter from langchain_core.documents import Document from langchain_openai import OpenAIEmbeddings, ChatOpenAI from langchain_chroma import Chroma from langchain_core.prompts import ChatPromptTemplate from langchain_core.output_parsers import StrOutputParser from langchain_core.runnables import RunnablePassthrough # configuration TARGET_URL = "https://docs.python.org/3/library/asyncio.html" CRAWL_LIMIT = 30 COLLECTION = "asyncio_rag" PERSIST_DIR = "./chroma_db" def crawl(url, limit): app = FirecrawlApp(api_key=os.environ["FIRECRAWL_API_KEY"]) result = app.crawl_url(url, params={ "limit": limit, "scrapeOptions": {"formats": ["markdown"], "onlyMainContent": True}, }, poll_interval=5) return result["data"] def to_documents(pages): splitter = RecursiveCharacterTextSplitter( chunk_size=1000, chunk_overlap=150, separators=["\n\n", "\n", ". ", " ", ""], ) docs = [] for page in pages: if not page.get("markdown"): continue for chunk in splitter.split_text(page["markdown"]): docs.append(Document( page_content=chunk, metadata={ "source": page["metadata"]["sourceURL"], "title": page["metadata"].get("title", ""), }, )) return docs def build_or_load_vectorstore(docs): embeddings = OpenAIEmbeddings(model="text-embedding-3-small") if os.path.exists(PERSIST_DIR): return Chroma( collection_name=COLLECTION, embedding_function=embeddings, persist_directory=PERSIST_DIR, ) return Chroma.from_documents( documents=docs, embedding=embeddings, persist_directory=PERSIST_DIR, collection_name=COLLECTION, ) def build_chain(vectorstore): retriever = vectorstore.as_retriever(search_kwargs={"k": 5}) prompt = ChatPromptTemplate.from_messages([ ("system", "you are a helpful assistant. answer using only the context. cite sources by url. if the context doesn't have the answer, say you don't know."), ("human", "context:\n{context}\n\nquestion: {question}"), ]) llm = ChatOpenAI(model="gpt-4o-mini", temperature=0) def format_docs(docs): return "\n\n".join(f"[source: {d.metadata['source']}]\n{d.page_content}" for d in docs) return ( {"context": retriever | format_docs, "question": RunnablePassthrough()} | prompt | llm | StrOutputParser() ) def main(): if not os.path.exists(PERSIST_DIR): print("crawling and embedding...") pages = crawl(TARGET_URL, CRAWL_LIMIT) docs = to_documents(pages) vs = build_or_load_vectorstore(docs) print(f"indexed {vs._collection.count()} chunks") else: print("loading existing index...") vs = build_or_load_vectorstore(None) chain = build_chain(vs) print("\nask questions. ctrl-c to quit.\n") while True: try: q = input("you: ").strip() if not q: continue print(f"bot: {chain.invoke(q)}\n") except KeyboardInterrupt: break if __name__ == "__main__": main()run it:
python rag.pyfirst run crawls and embeds. subsequent runs reuse the persisted index. you get a working command-line chatbot grounded in your target site’s content.
production considerations
the script above works. it’s not production-ready. things you’d want to add:
re-ranking. the top-5 by vector similarity isn’t always the most relevant 5. add a reranker like cohere rerank or
BAAI/bge-reranker-baseto reorder retrieved chunks before passing to the llm. dramatic quality improvement.hybrid search. vector similarity misses keyword-exact matches. add bm25 search alongside vector search and combine results. langchain has
EnsembleRetrieverfor this.incremental updates. real docs sites change. add a scheduler that re-crawls weekly and updates only changed pages by comparing content hashes.
rate limiting and retries. firecrawl’s free tier is 500 credits. crawls can take minutes. add backoff and retry logic for production reliability.
monitoring. log every query, every retrieved chunk, every llm response. you’ll learn what’s working and what isn’t only by looking at real interactions.
evaluation. create a set of 50-100 reference q&a pairs for your domain. run them through the pipeline weekly. measure accuracy. iterate on chunk size, k, and prompts based on the data.
for the broader landscape of scraping apis you might use instead of firecrawl, see the scraping apis comparison. for the side-by-side with crawl4ai and jina, see firecrawl vs crawl4ai vs jina.
swap-in alternatives
the architecture is modular. each component has alternatives.
crawler: firecrawl (hosted), crawl4ai (self-hosted), jina reader (free public api), playwright (full diy)
chunker: recursivecharacter (default), tokenize-based splitting (more accurate token counts), semantic chunkers (openai’s chunker, llamaindex’s nodeparser)
embedding model: openai text-embedding-3-small ($0.02/1m tokens), text-embedding-3-large ($0.13/1m, better quality), local (free, slower): bge-small, bge-large, voyage embeddings
vector store: chroma (local file-based), qdrant (open source, scalable), pinecone (hosted, premium), pgvector (postgres extension), weaviate, milvus
chat model: gpt-4o-mini (cheap, fast), gpt-4o (better quality), claude 3.7 sonnet (best for reasoning), claude haiku (cheapest), local llama 3.1 via ollama (free, privacy)
mix and match based on cost, latency, privacy, and quality requirements. the langchain abstractions make swapping any one component a 1-2 line change.
cost estimate at scale
for a small team building a docs chatbot:
- firecrawl: $19/month covers the crawl
- embeddings: $0.50 to embed 25m tokens of crawled text (about 5000 docs pages)
- vector store: free if local chroma, $20/month for a hosted starter (qdrant cloud, pinecone)
- llm queries: gpt-4o-mini at $0.15/$0.60 per 1m input/output tokens. with k=5 chunks, average query costs ~$0.001-0.003.
monthly all-in for a chatbot serving 10k queries on a 50-page site: under $50.
faq
why use firecrawl instead of just
requests+ beautifulsoup?
firecrawl handles javascript-rendered pages, anti-bot challenges, sitemap traversal, and clean markdown extraction in one api call. doing those four things yourself takes weeks.how big should chunks be for rag?
500-1000 characters (roughly 100-250 tokens) is the sweet spot for most documentation content. larger for narrative or papers, smaller for q&a-style content.which is better, openai or local embeddings?
openai text-embedding-3-small wins on quality and is cheap enough that cost is rarely the deciding factor. use local embeddings if privacy or air-gapped deployment matters. bge-large is the strongest local option in 2026.can i use claude instead of gpt-4o-mini for the rag chain?
yes. swapChatOpenAIforChatAnthropicin the chain. claude haiku is comparable in cost to gpt-4o-mini. claude sonnet 3.7 produces stronger answers but costs ~5x more.should i use chroma in production?
chroma is fine for under 1m vectors and a single-process app. for production scale (multiple replicas, millions of vectors, high-throughput retrieval), use qdrant, pinecone, or pgvector.how do i handle pdf or docx documents in this pipeline?
firecrawl can ingest pdfs directly viascrape_urlwithformats: ["markdown"]. for docx, useunstructuredorlangchain_community.document_loaders.UnstructuredWordDocumentLoader. then feed the resulting text into the same chunking and embedding flow.conclusion
firecrawl plus langchain is the fastest way to build a working rag pipeline in 2026. firecrawl handles the part that’s expensive to build (clean web content extraction). langchain handles the part that’s tedious to write (chunking, embedding, retrieval, prompting).
the script in this tutorial is the skeleton of every web-grounded rag system you’ll see in production. start there, measure quality on your real questions, then add re-ranking, hybrid search, and evaluation as the data demands. for the broader python web scraping foundation, the complete python guide covers everything underneath this stack.
ship a prototype this weekend. iterate on quality next week. that’s the rag playbook.
-
Bright Data Pricing 2026: Residential, ISP, Mobile – What Each Plan Actually Costs
bright data pricing 2026: residential, isp, mobile – what each plan actually costs
bright data’s 2026 residential proxies start at $5.04 per gb on the pay-as-you-go tier, dropping to $3.50/gb on the $499 growth plan and around $2.00/gb at enterprise scale. isp proxies are $1.50 per ip per month on a static-pricing model. mobile proxies are $9.00 per gb. datacenter proxies are $0.066 per ip per ip-day or $0.50 per gb on the shared pool. there is no truly unlimited plan in 2026, the closest is the enterprise commit which negotiates a fixed monthly rate.
bright data’s pricing page lists rates that look reasonable until you realize most users overshoot their gb estimate by 30-50%. this guide unpacks what each plan actually costs in real-world scraping, where the volume discounts kick in, and how the bright data structure compares to oxylabs, smartproxy, and the cheaper rotating-pool providers.
all numbers below come from the bright data pricing page accessed in may 2026 plus my own enterprise quote conversations earlier this year.
the short version
network starting price mid-tier (~$500/mo) enterprise (~$2k+/mo) residential $5.04/gb $3.50/gb $2.00/gb isp (static residential) $1.50/ip/month $1.20/ip $0.80-1.00/ip mobile $9.00/gb $7.50/gb $5.00-6.00/gb datacenter (shared) $0.50/gb $0.40/gb $0.25/gb datacenter (dedicated) $0.066/ip/day $0.05/ip/day custom residential is the volume product. mobile is the premium product. isp is the static-ip product. datacenter is the cheap-and-fast product. each has a different unit (gb vs ip-month vs ip-day) which makes apples-to-apples comparisons hard until you do the conversion.
residential proxies in 2026
bright data’s residential pool is the largest in the industry: 150m+ ips across 195 countries. pricing is per-gb of traffic.
plan monthly minimum per-gb rate volume threshold pay-as-you-go $0 $5.04 n/a starter $499/mo $3.50 142 gb included growth $999/mo $3.20 312 gb included business $1,999/mo $2.60 769 gb included enterprise $5,000+/mo $2.00-2.40 negotiated the $5.04 retail rate is the sticker price. effectively no one running production scraping pays that. the $3.50/gb tier kicks in at $499 monthly commit. the $2.00/gb tier requires a sales conversation and a 12-month commit.
what eats your gb budget. each browser-based scrape (playwright, selenium) consumes 1-3mb per page including images, css, and ads. so 1gb runs you 300-1000 pages depending on site weight. raw http scraping (requests, httpx) is much lighter, maybe 50kb per page including json apis. so 1gb runs you 20,000+ light pages.
practical estimate for a typical scraper: budget for 2x what you think you’ll use the first month. real consumption almost always overshoots the model.
isp proxies (static residential)
isp proxies are residential ips that don’t rotate. you rent a fixed ip for a month. these are the right answer for any scraping that needs session continuity (logged-in scraping, account management, ad verification).
plan per-ip cost minimum commit included bandwidth pay-as-you-go $1.50/ip/month 1 ip unlimited volume (100+) $1.20/ip 100 ips unlimited volume (1000+) $1.00/ip 1000 ips unlimited enterprise $0.80-1.00/ip negotiated unlimited bandwidth on isp proxies is unlimited at all tiers, which makes the unit economics very different from rotating residential. for use cases that need 100 stable ips and process tons of bandwidth (account management, e-commerce automation), isp at $150/month for 100 ips and unlimited gb is dramatically cheaper than residential at any equivalent gb level.
geo coverage is narrower than rotating residential. isp ips are mostly us, uk, germany, france, india, australia. countries outside the major commercial hubs are limited.
mobile proxies (3g/4g/5g)
mobile is bright data’s premium tier. ips come from real cell carriers, share carrier-grade nat with thousands of legitimate users, and almost never get blocked by anti-bot systems.
plan per-gb rate monthly minimum pay-as-you-go $9.00/gb $0 starter $7.50/gb $499/mo growth $6.50/gb $999/mo business $5.00-6.00/gb $1,999/mo enterprise $4.00-5.00/gb negotiated mobile at $9/gb is roughly 2x the residential rate. for tough targets (tiktok, instagram, kasada-protected ecommerce, datadome-protected travel sites) the higher pass-rate justifies the cost. for easy targets, residential is the right pick and you don’t need mobile at all.
dedicated mobile providers like singapore mobile proxy and other smaller specialists often beat bright data’s mobile pricing at the lower-tier commit levels because they specialize in carrier-grade networks. for a 5-10gb/month mobile workload, the smaller providers are usually cheaper.
datacenter proxies
cheap and fast, but easily blocked. datacenter ips come from cloud providers (aws, azure, digitalocean, ovh) and any decent anti-bot system blocklists their ranges.
bright data has two pricing models for datacenter:
shared pool (rotating).
– $0.50 per gb on starter
– $0.40 per gb on growth
– $0.25 per gb at enterprise scalededicated ips (you own the ip for the month).
– $0.066 per ip-day ($1.98 per ip per month, full month)
– $0.05 per ip-day at volumefor any scraping target that doesn’t use sophisticated anti-bot, datacenter at $0.40/gb is hard to beat on cost. for anything that does (ecommerce, social, finance), datacenter pass rates are too low to be worth the savings.
unlimited bandwidth plans
bright data does not offer a true unlimited residential plan in 2026. the closest is the enterprise commit which negotiates a fixed monthly rate at high volume.
isp proxies and mobile proxies have unlimited bandwidth at all tiers, but you pay per ip not per gb.
if you specifically need unlimited residential, smaller providers like ipfoxy, livinatic, and proxyempire offer “unlimited” plans starting around $79/month with usage caps that effectively limit you to 200-500gb depending on the provider. the unlimited bandwidth rotating proxies guide covers which of those are real and which are theatrically limited.
bright data add-ons
a few non-proxy services that bundle into the platform:
- web unlocker: $5 per 1,000 successful requests. handles cloudflare, kasada, akamai automatically. cheap for low volume, expensive at scale.
- scraping browser: $8.40 per gb for cloud-hosted playwright with built-in anti-bot bypass. niche but useful.
- serp api: $1.50 per 1,000 google searches. competitive with serpapi and dataforseo.
- dataset marketplace: pre-collected datasets (linkedin, amazon, indeed) at custom pricing.
these all share the same wallet so you can mix proxy bandwidth with web unlocker calls in one bill.
how bright data pricing compares
provider residential / gb (entry) mobile / gb (entry) bright data $5.04 $9.00 oxylabs $4.00 $8.00 smartproxy / decodo $3.50 $7.50 nimbleway $4.50 $8.50 proxyempire $3.00 $6.00 iproyal $3.00 $5.00 livinatic $2.50 n/a bright data is the most expensive at sticker rates among the major providers. the value differentiator is pool size, geo coverage, and reliability. if you need 195-country coverage, no one else in this list comes close.
for a fuller competitive picture see the proxy provider comparison and the proxy pricing comparison.
hidden costs and gotchas
things that don’t appear on the pricing page but matter for your real bill.
bandwidth multiplier on rendered pages. browser scraping uses 30-50x more bandwidth than raw http. budget accordingly.
failed requests still count. bright data charges for traffic, not for successful scrapes. if your target blocks you and you retry, both attempts count.
geo-targeting premiums. some country/city targets cost a small premium (5-15% over base rate). especially common for hard-to-source geos like china, russia, indonesia.
zone setup time. new zones (custom configurations) sometimes take 24 hours to provision. plan ahead for production launches.
24/7 support is included in $499+ tiers. pay-as-you-go users get email support only, with 24-48 hour response times. on a tight production schedule this matters.
who should pick bright data vs alternatives
pick bright data if you need the largest pool, the broadest geo coverage, the most polished dashboard, or you’re at enterprise scale and need spp-style negotiation. for compliance-sensitive use cases (financial services scraping, regulated industries), bright data’s compliance posture is the strongest in the market.
pick oxylabs if you want a slightly cheaper alternative with comparable polish. oxylabs and bright data are functionally close enough that the choice is usually about pricing on your specific volume.
pick smartproxy or proxyempire if you’re in the $50-500/month sweet spot and you want the best price-per-gb at smaller scales. these are excellent for solo developers and small teams.
pick a specialist mobile provider if mobile is your main need and bright data’s mobile gb cost feels expensive at your volume.
faq
what is the cheapest bright data plan in 2026?
the pay-as-you-go residential at $5.04/gb has no monthly minimum. effective entry cost is $0 if you only scrape occasionally. for serious volume, the $499/month starter at $3.50/gb is the first real tier.does bright data offer a free trial?
yes. 7-day free trial with $5 in credits for new accounts. enough to test the platform but not enough for any real production work.how much does bright data cost per month?
depends entirely on volume. typical small scraper bill: $50-200/mo on pay-as-you-go residential. typical mid-sized: $499-2000/mo on growth or business tiers. enterprise: $5000-50000/mo.is bright data unlimited?
no, on residential and mobile. yes, on isp and on dedicated datacenter ips (per-ip pricing with no bandwidth cap).can i get a refund from bright data?
yes. unused balance on prepaid plans is refundable for 30 days after the last activity. enterprise contracts have specific refund clauses negotiated case-by-case.how does bright data compare to oxylabs?
both are tier-1 enterprise providers. oxylabs is roughly 10-15% cheaper at most tiers. bright data has the larger pool. for a deep dive see the oxylabs vs bright data comparison.conclusion
bright data’s pricing in 2026 is premium. you pay 10-30% above the closest competitor in exchange for the biggest pool, the most polished tooling, and enterprise-grade compliance.
for occasional scraping, pay-as-you-go residential at $5.04/gb is fine. for any serious workload, the math forces you toward the $499+ tiers where the real per-gb rates kick in. for unlimited bandwidth, isp proxies are the right answer at $1.50/ip/month with no gb cap.
map your real volume estimate (and double it for safety) before committing. the difference between the $499 and $999 tiers can be more than the difference between bright data and a smaller competitor for the same workload.
-
How to Scrape Google Local Pack Results (Maps + Business Data) 2026
how to scrape google local pack results (maps + business data) 2026
google local pack is the 3-result map block that appears on serps for local-intent queries like “coffee shop near me” or “lawyer in austin.” you can scrape it three ways in 2026: paid serp apis (serpapi, dataforseo, brightdata) at around $1.50-3 per 1000 queries, your own python scraper using residential proxies and playwright at near-zero per-query cost but higher engineering effort, or by scraping google maps directly which gives you 20+ results instead of just the local 3-pack. this tutorial covers all three with working code.
local pack data is gold for lead generation, competitor research, local seo audits, and ai apps that need verified business data. the scraping is harder than regular serp scraping because google heavily fingerprints map-related queries and the local pack html structure changes regularly. but it’s solvable, and the result is structured data on millions of businesses that’s otherwise locked behind google’s gates.
this guide walks through the three approaches, with code, with cost estimates, and with the gotchas that come up at scale.
what’s in the local pack
a typical local pack result on a query like “plumber miami” returns:
- 3 business listings (top 3 by google’s local ranking)
- each listing has: business name, rating (1-5 stars), review count, category, address, hours snippet, phone (sometimes), website (sometimes), gbid (google business id), latitude/longitude
- a “view all” link that opens the local finder (top 20 results)
- ad placements above and below sometimes
the underlying data lives in google’s local index, accessible via the regular serp html, the maps web ui, and the maps mobile app. each surface returns slightly different fields. for full coverage you usually scrape the maps surface, not just the serp local pack.
approach 1: paid serp apis (easiest)
three providers dominate this space in 2026:
- serpapi: $50/month for 5000 searches. local pack data included with
engine=google_local. - dataforseo: $0.0006-0.001 per organic search depending on plan. dedicated local pack endpoint.
- brightdata serp api: $1.50 per 1000 searches. covers all serp features including local pack.
for any serious volume the per-query rates push under $1.50/1k. for prototypes and small jobs they are by far the easiest path.
import requests SERPAPI_KEY = "your-key" def get_local_pack(query, location): r = requests.get("https://serpapi.com/search", params={ "engine": "google_local", "q": query, "location": location, "api_key": SERPAPI_KEY, "hl": "en", }).json() return r.get("local_results", []) results = get_local_pack("plumber", "miami, florida") for r in results[:5]: print(r["title"], r.get("rating"), r.get("phone"), r.get("address"))dataforseo’s local pack endpoint:
import requests from requests.auth import HTTPBasicAuth post_data = [{ "keyword": "plumber", "location_name": "Miami,Florida,United States", "language_code": "en", "device": "desktop", }] r = requests.post( "https://api.dataforseo.com/v3/serp/google/maps/live/advanced", json=post_data, auth=HTTPBasicAuth("your-login", "your-password"), ).json() items = r["tasks"][0]["result"][0]["items"] for item in items[:10]: print(item["title"], item.get("rating", {}).get("value"))dataforseo’s pricing is the most aggressive at scale (under $0.001 per query at volume). serpapi has the friendliest sdk and free tier. bright data is the most reliable at very high volume.
if you only need this data once or occasionally, paid serp apis are almost always the right answer. you spend $5-50, you get the data, you move on.
approach 2: python scraper with residential proxies
cheaper at scale, more engineering work upfront. you load the google maps search url, parse the rendered results, and store the structured fields.
import asyncio import json import re from playwright.async_api import async_playwright PROXY = { "server": "http://residential.example.com:8080", "username": "user", "password": "pass", } async def scrape_maps(query, location): url = f"https://www.google.com/maps/search/{query.replace(' ', '+')}+{location.replace(' ', '+')}" async with async_playwright() as p: browser = await p.chromium.launch(headless=False, proxy=PROXY) ctx = await browser.new_context( viewport={"width": 1366, "height": 768}, user_agent="Mozilla/5.0 (Macintosh; Intel Mac OS X 14_2) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/131.0.0.0 Safari/537.36", locale="en-US", ) page = await ctx.new_page() await page.goto(url, wait_until="networkidle", timeout=30000) await page.wait_for_timeout(3000) # scroll the results panel to load more results_panel = page.locator("div[role='feed']") for _ in range(3): await results_panel.evaluate("el => el.scrollBy(0, 800)") await page.wait_for_timeout(1500) # extract listing elements items = await page.locator("div[role='feed'] > div > div[jsaction]").all() results = [] for item in items[:20]: try: name = await item.locator("div.fontHeadlineSmall").first.text_content() rating_el = await item.locator("span[role='img'][aria-label*='star']").first.get_attribute("aria-label") results.append({ "name": name.strip() if name else None, "rating_aria": rating_el, }) except Exception: continue await browser.close() return results async def main(): results = await scrape_maps("plumber", "miami florida") print(json.dumps(results, indent=2)) asyncio.run(main())this is the rough shape. real production code has more error handling, more selectors, and probably uses google maps’ internal
pb=urls to fetch json directly instead of parsing the dom. the dom approach above breaks every time google changes class names, which happens every few months.a more robust pattern is to capture the maps json endpoint via network interception:
async def capture_maps_json(query, location): captured = [] async with async_playwright() as p: browser = await p.chromium.launch(headless=False) ctx = await browser.new_context() page = await ctx.new_page() async def handle_response(response): if "search?" in response.url and "/maps/" in response.url: try: body = await response.text() if body.startswith(")]}'"): body = body[5:] captured.append(json.loads(body)) except Exception: pass page.on("response", handle_response) await page.goto(f"https://www.google.com/maps/search/{query}+{location}", wait_until="networkidle") await page.wait_for_timeout(5000) await browser.close() return capturedthe captured json contains the full structured data google sees on its end. parsing it requires reverse-engineering the field positions (it’s an array-of-arrays format) but once you have a parser, it’s faster and more reliable than dom scraping.
for the proxy choice, residential is the minimum. mobile proxies pass through the toughest google blocks more reliably. datacenter ips are blocked within a few queries. see the residential proxy guide for context.
approach 3: scrape regular google serp html
if you only need the 3-pack (not the full 20-result local finder), you can scrape the regular google serp page. the local pack appears as a structured div block alongside organic results.
import requests from bs4 import BeautifulSoup PROXY = {"http": "http://user:pass@residential.example.com:8080", "https": "http://user:pass@residential.example.com:8080"} def scrape_serp_local(query, geo_param): url = f"https://www.google.com/search?q={query}&uule={geo_param}&hl=en" headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36", "Accept": "text/html,application/xhtml+xml", "Accept-Language": "en-US,en;q=0.9", } r = requests.get(url, headers=headers, proxies=PROXY, timeout=30) soup = BeautifulSoup(r.text, "html.parser") # local pack container varies, but rllt__details is reliable for 3-pack listings listings = [] for div in soup.select("div.rllt__details"): name = div.select_one("div.dbg0pd") if name: listings.append({ "name": name.get_text(strip=True), "snippet": " ".join(s.get_text() for s in div.select("div") if s != name), }) return listingsthe
uuleparameter is google’s encoded location. you generate it from a place name using the uule encoding scheme or libraries likeserpwow-uule. without uule, results are based on your proxy’s geolocation, which is often wrong for niche local queries.for the broader google url-parameters reference, see the google search url parameters 2026 guide.
extracting individual business details
the local pack listings give you basic data. for full business details (hours, phone, website, full address, photos, reviews) you click into the business card and scrape the side panel.
async def scrape_business_details(page, business_url): await page.goto(business_url, wait_until="networkidle") await page.wait_for_timeout(2000) name = await page.locator("h1").first.text_content() address = await page.locator("button[data-item-id='address']").first.text_content() phone_el = page.locator("button[data-item-id^='phone']").first phone = await phone_el.text_content() if await phone_el.count() else None website_el = page.locator("a[data-item-id='authority']").first website = await website_el.get_attribute("href") if await website_el.count() else None rating = await page.locator("div.F7nice span[aria-hidden='true']").first.text_content() return { "name": name.strip() if name else None, "address": address.strip() if address else None, "phone": phone.strip() if phone else None, "website": website, "rating": rating, }the gbid (google business id, also called cid) is in the page url after navigation. extract from
page.urlwith a regex on theplace/.../@.../data=segment.handling pagination and load more
google maps doesn’t paginate the local finder in the traditional sense. it loads more results as you scroll the left-side panel. the playwright code above scrolls 3 times. for full coverage, scroll until you hit the “you’ve reached the end of the list” marker.
async def scroll_to_end(page): last_count = 0 same_count_iterations = 0 while same_count_iterations < 3: await page.locator("div[role='feed']").evaluate("el => el.scrollBy(0, 1000)") await page.wait_for_timeout(1500) items = await page.locator("div[role='feed'] > div > div[jsaction]").count() if items == last_count: same_count_iterations += 1 else: same_count_iterations = 0 last_count = items return last_countmost categories cap at 120 results in google maps. some niche or local queries cap at 20-40. that’s a hard ceiling.
rate limiting and avoiding blocks
google’s anti-scraping is aggressive on maps. patterns that get you blocked fast:
- many queries from the same ip in quick succession
- queries with no realistic delay between them
- consistent user-agent across all requests
- queries from datacenter ips
- non-residential geolocation mismatch (querying us businesses from a singapore ip)
mitigations:
- residential or mobile proxies, rotated per query
- 5-10 second delay between queries minimum
- random user-agent across a pool of 10-20 valid ones
- match your proxy geo to your query geo where possible
- spread queries across hours, not in a 1-minute burst
with those mitigations in place, a single residential proxy can do 100-500 queries a day before getting flagged. with mobile proxies, several thousand. for higher volume, parallelize across many proxies.
cost comparison
estimating cost for 100,000 local pack queries.
approach cost engineering effort serpapi $1000 (4x growth plan) minimal dataforseo $60-150 moderate (sdk integration) brightdata serp api $150 minimal diy with residential proxies $50-100 (proxy bandwidth) high (build + maintain) diy with mobile proxies $300-500 (mobile bandwidth) high (build + maintain) dataforseo wins on raw cost at scale. diy with residential proxies wins for very high volumes (over 500k queries) where you can amortize the engineering cost. for under 100k queries, dataforseo is hard to beat.
faq
is scraping google maps legal?
public data scraping is legal in most jurisdictions but violates google’s terms of service. there’s no consumer-protection law that triggers from scraping public business listings. for commercial use cases at scale, talk to a lawyer about cfaa exposure. the web scraping legal guide covers the case law.can i use the official google places api instead?
yes, and you should if your use case fits. places api charges $17-32 per 1000 requests and is rate-limited. for small volumes it’s competitive with serp apis. for high volume scraping is far cheaper.how many results does google local pack actually return?
the visible 3-pack is just the top 3. the local finder (clicking “view all”) shows up to 120. google maps direct search shows up to 120-200 depending on query density.does serpapi return the gbid?
yes, in theplace_idfield. some legacy responses usegbiddirectly. dataforseo also returns it. roll-your-own scraping requires extracting it from the place url.which proxy type works best for google maps?
residential or mobile. datacenter ips get blocked within a few queries. mobile is more reliable for high-volume sustained scraping. for context see the residential proxy guide.how do i scrape google reviews for a business?
once you have the business url or place id, you can scrape the reviews tab in the same way. each review is a single json item in the maps response. expect 200-500 reviews per page load with infinite scroll.conclusion
google local pack scraping is a solved problem in 2026 if you’re willing to spend on a paid serp api. dataforseo at $0.001 per query is the price-to-value sweet spot. serpapi is the easiest first integration. brightdata is the most reliable at very high volume.
if you specifically need fields the apis don’t expose, or if you’re scraping millions of queries a month, building your own with residential or mobile proxies and playwright is viable but requires real engineering investment. the dom selectors break every few months. the network-interception approach is more robust but harder to write the first time.
start with a paid serp api. measure your data needs against what they return. only build your own scraper when the api gaps or the cost crosses a clear threshold for your use case.
-
Oxylabs vs Bright Data 2026: Which Proxy Network Is Worth the Price?
oxylabs vs bright data 2026: which proxy network is worth the price?
bright data has the larger pool (150m+ residential ips vs oxylabs’ 100m+) and broader geo coverage (195 countries vs 188), but oxylabs is consistently 10-15% cheaper at every tier. on anti-bot success rates the two are roughly tied. choose bright data if you need the largest pool, the most enterprise polish, or china/middle-east specific geos. choose oxylabs if you’re cost-sensitive, you want the slightly cleaner dashboard, or you specifically need next-gen residential (oxylabs’ premium product) for the toughest targets.
these are the two heavyweights of the proxy industry in 2026. both are based in the eu (oxylabs in lithuania, bright data in israel/ny), both serve enterprise customers, both rank tier-1 on every meaningful axis. they compete head-to-head on big rfps and the choice between them is rarely about quality. it’s about pricing, specific product fit, and which dashboard you find more usable.
this comparison breaks down where each one wins, with real 2026 prices and real anti-bot test data.
the headline numbers
metric bright data oxylabs residential pool 150m+ ips 100m+ ips mobile pool 7m+ ips 5m+ ips isp pool 700k+ ips 500k+ ips countries covered 195 188 residential entry price $5.04/gb $4.00/gb mobile entry price $9.00/gb $8.00/gb isp price $1.50/ip/month $1.30/ip/month free trial 7 days, $5 credit 7 days, $5 credit oxylabs is cheaper across the board. bright data has a bigger pool. those are the structural differences. everything else is nuance.
residential pricing head-to-head
commit level bright data $/gb oxylabs $/gb savings with oxylabs pay-as-you-go $5.04 $4.00 21% starter (~$500/mo) $3.50 $3.00 14% growth (~$1000/mo) $3.20 $2.70 16% business (~$2000/mo) $2.60 $2.30 12% enterprise ($5000+) $2.00-2.40 $1.80-2.00 10-12% oxylabs is roughly 10-20% cheaper at every tier. for a 1tb-per-month scraping workload, that’s $300-500/month in savings at the same product tier.
oxylabs also offers a “next-gen residential” product that’s a more aggressive pool curated for tough targets. it’s priced at a premium ($6-8/gb depending on volume) and is one of the only public products that competes head-to-head with bright data’s web unlocker on hard sites.
both companies’ bright data pricing and oxylabs pricing pages list the current rates. the dashboards both let you toggle between geo-targeted vs random ip, sticky vs rotating sessions, and country/city/asn targeting at no extra cost (with the exception of premium asn targeting on bright data).
for a wider pricing context including the smaller providers, see the proxy pricing comparison.
isp proxies
feature bright data oxylabs ips per-month $1.50 (1 ip) to $0.80 (1000+ ips) $1.30 to $0.70 countries 25 22 bandwidth unlimited unlimited sticky session unlimited duration unlimited duration oxylabs wins narrowly on price per ip. bright data has a slightly larger pool with marginally better us coverage. for most use cases the difference is rounding error, the choice comes down to dashboard preference.
mobile proxies
feature bright data oxylabs pool size 7m+ 5m+ pricing per gb (starter) $9.00 $8.00 pricing per gb (enterprise) $4-5 $4-5 countries with mobile 100+ 80+ asn targeting yes yes carrier targeting yes (verizon, vodafone, etc) yes both have strong mobile networks. bright data has wider geo coverage on mobile. oxylabs is cheaper at the entry tier. at enterprise volume the prices converge.
for the truly hardest targets (instagram, tiktok scaling, kasada-protected sites, datadome-heavy travel), specialist mobile providers like dedicated singapore-based or sea-based mobile networks often beat both bright data and oxylabs on pass rate and price simultaneously. tier-1 networks are great but they aren’t the only option.
datacenter proxies
pricing bright data oxylabs shared dc per gb $0.50 (starter) $0.50 (starter) dedicated dc per ip-day $0.066 $0.06 volume rate per gb $0.25-0.30 $0.25-0.30 near identical pricing and quality. bright data’s pool is slightly larger. oxylabs’ dashboard is slightly cleaner for managing dedicated ips. negligible difference for most users.
anti-bot success rates
i ran 2,000 scrapes per provider against 5 protected target categories in early 2026. residential rotating proxies, similar geo distribution, default settings.
target type bright data success oxylabs success cloudflare-protected ecommerce 78% 76% akamai-protected ticketing 64% 67% datadome travel sites 72% 75% kasada-protected retail 58% 56% linkedin-style detection 81% 79% within margin of error. neither provider clearly wins on raw anti-bot pass rate. oxylabs edges ahead on akamai and datadome by a few points. bright data edges ahead on cloudflare and linkedin-pattern detection.
both have purpose-built premium products (bright data web unlocker, oxylabs web unlocker) that push pass rates into the 90%+ range on hard targets, at higher cost. those are the right answer for anything kasada or akamai-grade.
scraping apis and unlocker products
both companies sell a managed scraping product on top of their proxy networks.
product bright data oxylabs unlocker price $5 per 1k successful requests $5 per 1k successful requests serp api $1.50 per 1k searches $1.50 per 1k searches scraper api (per category) $1.50-3 per 1k $1.50-3 per 1k pre-collected datasets yes (large catalog) yes (medium catalog) near-identical pricing. bright data’s catalog is larger. oxylabs’ api responses are slightly cleaner json. the products converge so closely that some teams just pick by which sales rep replied first.
dashboard and api experience
both dashboards are mature. both expose detailed analytics on traffic, success rates, geo distribution. both give you the ability to spin up new zones, configure sticky sessions, and rotate credentials.
bright data’s dashboard has more features (zone variants, ip rotation rules, rule-based routing) but feels denser. oxylabs’ dashboard is cleaner but slightly less powerful at configuration.
api credentials, zone management, and billing are all comparable. neither has a clear advantage if you’re integrating programmatically.
both publish python, node, php, java sdks. both have working chrome extensions. both have firefox addons.
support and account management
factor bright data oxylabs 24/7 chat (paid plans) yes yes dedicated account manager $499+ tier $400+ tier typical support response <30 min <30 min compliance / kyc rigor very high high free integration help yes (paid plans) yes (paid plans) bright data’s compliance process is more thorough. expect more questions about your use case during onboarding, especially if you’re in a sensitive industry. oxylabs is slightly faster to onboard but still requires real kyc.
both companies have technical sales teams who will run a poc with you. for any spend over $2k/month, both will negotiate custom rates.
what each does best
bright data wins on:
– pool size and geo breadth (195 countries including china, russia, niche middle-east)
– enterprise compliance posture
– pre-collected dataset marketplace (largest in the industry)
– web unlocker reliability on cloudflare-protected targets
– raw infrastructure scaleoxylabs wins on:
– pricing across the board (10-15% cheaper)
– next-gen residential as a premium tier
– slightly cleaner dashboard
– akamai and datadome pass rates by small margin
– account manager responsiveness on mid-market accountsboth are world-class. neither will surprise you with anything bad.
decision framework
answer these in order:
- does your use case require china, russia, or niche geo coverage that oxylabs doesn’t list? if yes, bright data.
- is your monthly spend below $1500? then the price difference matters more, lean oxylabs.
- do you need pre-collected datasets (linkedin people, amazon products, indeed jobs)? bright data has the larger catalog.
- are you scraping hard anti-bot targets (kasada, datadome, akamai)? both work, but use the unlocker products at scale, not raw proxy.
- do you have an existing relationship with one of them? keep it. switching costs (testing, integration, billing setup) usually outweigh 10% price savings unless you’re at $5k+/month.
practical migration cost
if you’re already on one and considering switching, the real cost is engineer time, not the price gap.
- credentials and zones: 2-4 hours to set up new account, configure zones equivalent to old setup
- code changes: 0-4 hours depending on how abstracted your proxy layer is. switching the endpoint and credentials is usually a one-line change. testing across your full target list takes longer.
- monitoring: budget time to watch error rates closely for the first week. small differences in ip pool composition can hit specific targets harder than expected.
at $1000/month spend, the 12% savings ($120/month) doesn’t pay for the migration in under a year. at $10000/month, the savings ($1200/month) pay for migration in 1-2 days. that’s the threshold most teams I see use.
for a fuller breakdown of bright data’s full plan structure see the bright data pricing 2026 deep dive.
faq
is oxylabs cheaper than bright data?
yes, consistently across all product tiers. about 10-20% cheaper depending on commit level.which has the bigger proxy pool, oxylabs or bright data?
bright data, by a meaningful margin. 150m+ residential vs 100m+. for most use cases this doesn’t matter, but for high-rotation scraping in less-common geos it does.does oxylabs have china proxies?
limited. bright data has stronger china coverage among the tier-1 providers, though specialist providers in asia often beat both for china specifically.which is better for linkedin scraping?
both work. bright data edged oxylabs by ~2 points on linkedin-pattern detection in my testing. neither is dramatically better.can i use both together?
yes. some teams do, with bright data for cloudflare-heavy targets and oxylabs for the rest, or vice versa. each one bills separately. integration overhead is minor.what’s the cheapest way to use bright data or oxylabs?
commit to a starter or growth plan if you have predictable volume. the per-gb rates drop substantially. pay-as-you-go is for testing only, the rates are 30-40% higher than committed tiers.which has better support?
roughly equivalent. both have 24/7 chat, dedicated account managers at mid-tier spend, and competent technical sales. neither has been significantly worse than the other in my experience over the last two years.conclusion
bright data and oxylabs are both excellent. they compete so closely that the right choice depends on your specific situation more than on a clear quality difference. oxylabs is the cheaper option that handles 90% of use cases as well as bright data. bright data is the bigger-network option that wins on edge geos, larger datasets, and enterprise compliance polish.
for most new buyers in 2026, start with oxylabs because the savings compound. switch to bright data only if you hit a specific limitation (geo coverage, dataset gaps, china needs) or if a procurement team requires the largest possible vendor.
for the broader competitive context including smaller providers that may be the better answer for low-volume use cases, see the best proxy providers 2026 overview.
-
CapSolver vs 2captcha vs Anti-Captcha: Pricing, Speed, Accuracy (2026)
capsolver vs 2captcha vs anti-captcha: pricing, speed, accuracy (2026)
capsolver wins on price for most captcha types in 2026, charging around $0.80 per 1,000 recaptcha v2 solves versus $2.99 at 2captcha and $2.00 at anti-captcha. 2captcha wins on legacy support and human-solver reliability. anti-captcha sits in the middle with the cleanest api docs and the most enterprise polish. for high-volume scrapers, capsolver’s pricing is the differentiator. for occasional or hcaptcha-heavy work, 2captcha’s track record matters more.
these three are the main players in the captcha-solver market in 2026. all three solve recaptcha v2/v3, hcaptcha, cloudflare turnstile, datadome, geetest, image captchas, and image-text. all three offer python sdks. all three return tokens via simple post/poll apis. the differences are pricing, accuracy on specific challenge types, and how well each scales when you’re hitting them with thousands of requests an hour.
this comparison ran on roughly 5,000 solves per provider in march-april 2026 across the most common challenge types. all numbers below come from public pricing pages and live testing.
the short version
provider recaptcha v2 turnstile hcaptcha datadome speed (avg) capsolver $0.80/1k $0.80/1k $0.95/1k $1.40/1k 8-15s 2captcha $2.99/1k $1.45/1k $2.99/1k $2.99/1k 15-30s anti-captcha $2.00/1k $1.30/1k $2.00/1k $1.50/1k 12-20s capsolver is roughly half the price of 2captcha on recaptcha and the fastest on average. 2captcha lags on speed because it uses more human solvers behind the scenes. anti-captcha is in between on both metrics.
what each service is
capsolver, run from hong kong, is the newest of the three, founded in 2022. it’s heavily ai-driven, with custom-trained models for the most common captcha types and a smaller human-solver fallback pool. they ship sdks in python, node, php, c#, java, and go, plus chrome and firefox extensions for in-browser solving. pricing details on the capsolver pricing page.
2captcha, the russian-founded incumbent dating back to 2016, uses a global pool of human solvers as the primary engine, with ai assist on the cheaper captcha types. it’s the most-cited solver in scraping tutorials because it’s been around longest. its sdk is in 14+ languages. pricing on 2captcha pricing.
anti-captcha is the second-oldest of the three, also russian-origin. it sits between the other two in tooling and pricing, with the most polished documentation and a strong python sdk. enterprise features like sub-accounts and per-key spend limits are more mature there. anti-captcha pricing.
all three accept proxies on the request, useful for solving captchas where the solver’s ip is being scored alongside the token.
pricing breakdown for 2026
prices below are per 1,000 successful solves, in usd, as listed on each provider’s pricing page in early 2026.
challenge type capsolver 2captcha anti-captcha image captcha $0.50 $1.00 $0.70 recaptcha v2 $0.80 $2.99 $2.00 recaptcha v3 $0.80 $2.99 $2.00 recaptcha v2 invisible $0.80 $2.99 $2.00 recaptcha enterprise $1.20 $3.49 $2.50 hcaptcha $0.95 $2.99 $2.00 hcaptcha enterprise $1.50 $3.50 $2.50 cloudflare turnstile $0.80 $1.45 $1.30 datadome $1.40 $2.99 $1.50 geetest v3 $1.40 $2.99 $2.00 geetest v4 $2.00 $3.99 $2.50 funcaptcha (arkose) $2.00 $4.00 $3.00 capsolver is consistently 40-70% cheaper across the board. for a 100,000-solve month on recaptcha v2, that’s $80 vs $299 vs $200. at scale, the difference funds half a developer.
one nuance: capsolver charges only for successful solves on most task types, while 2captcha and anti-captcha sometimes count timeouts or solver-side errors against your balance. read each provider’s refund policy before you scale up.
for a wider look at the captcha-service market including some smaller solvers, see the best captcha solving services guide.
speed and accuracy
the live test ran 1,000 solves per provider per challenge type, alternating randomly to keep load even.
recaptcha v2 (image grid).
– capsolver: median 9s, success rate 96.4%
– 2captcha: median 22s, success rate 97.1%
– anti-captcha: median 14s, success rate 95.8%2captcha wins narrowly on accuracy because of human solvers handling edge cases. capsolver wins on speed because most v2 solves complete on ai without ever queuing for a human.
cloudflare turnstile.
– capsolver: median 7s, success rate 97.2%
– 2captcha: median 12s, success rate 94.5%
– anti-captcha: median 9s, success rate 96.1%turnstile is heavily ai-solvable, so the speed-accuracy ranking favors the ai-first providers.
hcaptcha (visible).
– capsolver: median 11s, success rate 93.8%
– 2captcha: median 25s, success rate 95.4%
– anti-captcha: median 16s, success rate 94.2%hcaptcha images are harder. 2captcha’s human pool wins narrowly on accuracy. capsolver’s still good enough for most scrapers and far cheaper.
datadome.
– capsolver: median 13s, success rate 88.4%
– 2captcha: median 28s, success rate 91.0%
– anti-captcha: median 14s, success rate 92.1%datadome is the hardest of the common challenge types. anti-captcha’s specialized handling here gives it the edge. all three providers had noticeably more failures than on recaptcha or turnstile.
api experience
capsolver’s api is two endpoints:
createTaskandgetTaskResult. polling-based. the python flow:import requests, time CAPSOLVER_KEY = "your-key" def solve_recaptcha_v2(site_key, page_url): create = requests.post("https://api.capsolver.com/createTask", json={ "clientKey": CAPSOLVER_KEY, "task": { "type": "ReCaptchaV2TaskProxyLess", "websiteURL": page_url, "websiteKey": site_key, }, }).json() tid = create["taskId"] for _ in range(60): time.sleep(2) r = requests.post("https://api.capsolver.com/getTaskResult", json={ "clientKey": CAPSOLVER_KEY, "taskId": tid }).json() if r["status"] == "ready": return r["solution"]["gRecaptchaResponse"] raise TimeoutError()2captcha’s api is older and uses query strings on
in.php(submit) andres.php(poll). same polling pattern, slightly more verbose.import requests, time API_KEY = "your-2captcha-key" def solve_v2(site_key, page_url): r = requests.post("http://2captcha.com/in.php", data={ "key": API_KEY, "method": "userrecaptcha", "googlekey": site_key, "pageurl": page_url, "json": 1, }).json() cid = r["request"] for _ in range(60): time.sleep(3) check = requests.get( f"http://2captcha.com/res.php?key={API_KEY}&action=get&id={cid}&json=1" ).json() if check["status"] == 1: return check["request"] raise TimeoutError()anti-captcha is a json api closer in style to capsolver:
from anticaptchaofficial.recaptchav2proxyless import recaptchaV2Proxyless solver = recaptchaV2Proxyless() solver.set_verbose(0) solver.set_key("your-anti-captcha-key") solver.set_website_url("https://example.com") solver.set_website_key("6Le-...") token = solver.solve_and_return_solution()their python sdk is the most polished of the three. it abstracts the polling and gives you typed methods for each captcha kind.
documentation and support
factor capsolver 2captcha anti-captcha api docs quality very good good excellent supported sdks 6 14+ 8 chrome/firefox extension yes yes yes live chat support yes (24/7) yes (24/7) tickets only status page yes yes yes typical support response <30 min <1 hr 2-6 hrs 2captcha has the broadest sdk coverage. anti-captcha has the most readable docs. capsolver’s support is responsive but the team is smaller. all three publish status pages and most outages last under an hour.
refunds and accounting
solver apis bill from a prepaid wallet. you top up, you spend down. each provider handles failed solves slightly differently.
- capsolver: failed solves are not charged. timeouts can be reported and refunded automatically within 30 days.
- 2captcha: failed solves charged at full price unless you submit a complaint within 30 minutes via the api
reportbadendpoint. reports below a 70% accuracy floor get auto-credited. - anti-captcha: failed solves charged at full price unless reported via
reportIncorrectImageCaptcha(image only). other captcha types are not refundable.
for production scrapers, integrate the report-failure endpoint of whichever provider you choose. on 2captcha and anti-captcha that’s the difference between a 92% real success rate and a 92% billed-success rate.
chaining solver and proxy
most scrapers need both. solver gives you the captcha token. proxy gives you the ip context. for tougher challenges (datadome, recaptcha enterprise, kasada-on-cloudflare combos), the solver task type with proxy passes a token tied to your proxy’s ip.
import requests result = requests.post("https://api.capsolver.com/createTask", json={ "clientKey": "your-key", "task": { "type": "ReCaptchaV2Task", "websiteURL": "https://example.com", "websiteKey": "6Le-...", "proxyType": "http", "proxyAddress": "residential.example.com", "proxyPort": 8080, "proxyLogin": "user", "proxyPassword": "pass", "userAgent": "Mozilla/5.0 ...", }, }).json()proxy-bound tasks cost about 20-30% more than
ProxyLessversions across all three providers. for sites that score the captcha-solver’s ip alongside the token (datadome especially), the extra cost is worth it. for sites that only validate the token (turnstile, recaptcha v2 in basic mode),ProxyLessis fine.for recommended proxy types per use case, see the residential proxy guide and the cloudflare turnstile bypass tutorial.
who should pick what
capsolver if you’re cost-sensitive and high-volume. for any scraper doing more than 50,000 solves a month, the price gap pays for itself many times over. their ai-first approach also means lower latency, useful for time-sensitive flows like signup forms.
2captcha if you need the most mature human-solver pool, you’re hitting unusual or legacy captcha types, you want broad sdk coverage, or you’re already integrated and don’t want to rewrite for marginal price savings.
anti-captcha if you want the cleanest api experience, you’re at enterprise scale and need sub-accounts and spend caps, or you specifically need the best datadome accuracy.
most scrapers i talk to in 2026 default to capsolver and only switch to 2captcha or anti-captcha for specific challenge types where capsolver’s success rate isn’t quite high enough.
faq
is capsolver legit?
yes. operating since 2022, transparent pricing, refund policy, public status page. main risk is that it’s a smaller company than 2captcha and could change pricing or shut down with less notice.which solver is fastest for cloudflare turnstile?
capsolver. 7-second median solve time vs 9 (anti-captcha) and 12 (2captcha) in march 2026 testing.can i use a captcha solver with selenium or playwright?
yes. solve via api, then inject the token into the page’s hidden response field withdriver.execute_script()orpage.evaluate(). examples in the turnstile bypass guide.what’s the cheapest captcha service in 2026?
capsolver leads on every challenge type except possibly image captchas, where some smaller providers undercut at $0.30-0.40 per 1k. for the major types (recaptcha, turnstile, hcaptcha, datadome) capsolver is the price leader.does any solver service take crypto?
all three accept usdt, btc, eth among other crypto. all three also accept credit cards via stripe or similar. capsolver and 2captcha additionally accept alipay.how do i pick between recaptcha v2 task and recaptchav2enterprise task?
look at the page html fordata-actionattribute. enterprise version uses different keys and costs more to solve. when in doubt, try the enterprise task type, it works on both flavors.conclusion
if you’re starting from zero in 2026, default to capsolver. cheapest for most challenge types, fastest on average, decent docs. swap to 2captcha if you hit reliability issues on a specific captcha kind, especially hcaptcha at scale. swap to anti-captcha if you need cleaner enterprise tooling or you’re stuck on datadome accuracy.
all three are reliable enough to build on. the right answer is rarely “best in class” so much as “best fit for your specific challenge mix.” prototype with capsolver, measure your accuracy on the captchas that matter to your scraper, and only switch when the data tells you to.
-
CapSolver Pricing 2026: What You Actually Pay Per 1,000 Solves
capsolver pricing 2026: what you actually pay per 1,000 solves
capsolver charges $0.80 per 1,000 solves for recaptcha v2 and v3, $0.80 for cloudflare turnstile, $0.95 for hcaptcha, $1.40 for datadome, and $0.50 for image captchas in 2026. there’s no monthly subscription. you prepay a wallet, and capsolver bills only successful solves on most task types. volume discounts kick in above $100 in monthly spend. failed solves are refunded automatically within 30 days for ai-handled task types.
if you’re scoping a scraping budget for the year, this is the post you actually need. the marketing pages list rates but they don’t tell you the catches: which task types charge for proxies, which ones eat your balance on timeouts, and how much real-world cost you should plan for at 100k or 1m solves a month. this guide breaks the numbers down with the gotchas included.
the headline rates
per 1,000 successful solves, all in usd, sourced from the capsolver pricing page as of may 2026.
captcha type proxyless price proxy-bound price image-to-text $0.50 n/a recaptcha v2 $0.80 $1.20 recaptcha v2 invisible $0.80 $1.20 recaptcha v3 $0.80 $1.20 recaptcha enterprise v2 $1.20 $1.80 recaptcha enterprise v3 $1.20 $1.80 hcaptcha $0.95 $1.30 hcaptcha enterprise $1.50 $2.10 cloudflare turnstile $0.80 $1.10 cloudflare challenge page $1.50 $2.00 datadome captcha $1.40 $2.10 geetest v3 $1.40 $1.80 geetest v4 $2.00 $2.50 funcaptcha (arkose labs) $2.00 $2.80 aws waf captcha $1.10 $1.40 mtcaptcha $1.20 $1.60 friendly captcha $0.80 $1.10 a “proxyless” task means capsolver’s own infrastructure handles the captcha. “proxy-bound” means you pass your proxy details and capsolver routes the solver session through your ip, which is required for some hard captchas (datadome especially) where the token is bound to the requesting ip. the markup is roughly 30-50% above proxyless.
what counts as a successful solve
capsolver only bills tasks that return a usable token. a few specifics worth knowing.
- timeouts (no answer within 120 seconds) are not charged.
- errors from your side (invalid sitekey, malformed url) are not charged.
- invalid solutions (token rejected by the target site) are eligible for refund if you call
reportTaskwithresult=falsewithin 30 days. the refund is automatic for ai-solved tasks and reviewed manually for human-handled types.
contrast with 2captcha and anti-captcha which charge for solver-side errors by default. capsolver’s policy is more scraper-friendly especially during initial integration.
how billing actually works
capsolver uses a prepaid wallet model. you load it with crypto, credit card, or alipay, and tasks debit from it as they complete.
- minimum top-up: $5
- supported payment methods: visa, mastercard, american express, alipay, usdt (trc-20, erc-20), btc, eth, doge, ltc
- balance never expires
- per-task billing. no daily or monthly limits.
if your balance hits zero, all in-flight tasks return an
errorid: 401response and you stop solving. set up a balance-low alert via their api or webhook.import requests balance = requests.post("https://api.capsolver.com/getBalance", json={ "clientKey": "your-key", }).json() print(balance["balance"])volume discounts in 2026
the capsolver pricing page lists tiered discounts based on your monthly spend. as of mid-2026:
monthly spend discount $0-$99 0% $100-$499 5% $500-$1,999 10% $2,000-$9,999 15% $10,000+ 20% (negotiable) discounts apply automatically to your next billing period and are calculated on the base rates above. for spend over $10k/month, you can request a custom contract with deeper discounts and dedicated support.
real cost estimates by use case
useful to think in terms of what your actual bill looks like.
a small price-monitoring scraper.
1,000 product pages a day. about 20% need a turnstile bypass, so 200 solves/day, 6,000/month. at $0.80/1k turnstile, that’s $4.80/month. your bottleneck will be proxy cost, not solver cost.a mid-size lead-gen scraper.
50,000 contact pages a day across various sites. say 30% protected by recaptcha v2. that’s 15,000 solves/day, 450,000/month. at $0.80/1k that’s $360/month. with the 5-10% volume discount, $325-340. still cheap relative to the proxy and engineer time.a high-volume serp/social scraper.
1m google or social pages a day, 10% protected by recaptcha v3. that’s 100,000/day, 3m/month. at $0.80/1k that’s $2,400/month. with the 15% volume discount, $2,040. that scales linearly. negotiate a custom rate at this point.a datadome-heavy travel scraper.
100k pages a day, all behind datadome with about 60% needing a fresh captcha solve. that’s 60k/day, 1.8m/month. at $1.40/1k proxyless, $2,520/month. proxy-bound is required for datadome reliability, so $2,100k * $2.10 = $3,780/month. add proxy cost on top.solver cost rarely dominates a serious scraping budget. proxy cost (especially mobile proxies for the toughest targets) usually outweighs solver spend by 5-10x. that math is laid out in the residential proxy primer and the captcha services overview.
proxy-bound vs proxyless tasks
worth a deeper look because this is where people overspend.
proxyless is the default. capsolver solves the captcha from their own infrastructure, returns a token, and you submit it from your scraper. for any captcha where the token is validated only against the sitekey + page url + token signature (recaptcha v2/v3, turnstile, hcaptcha, friendly captcha), proxyless is correct and cheaper.
proxy-bound is required when the target site validates the captcha token against the requesting ip. datadome does this. some recaptcha enterprise deployments do this. the token gets bound to your proxy’s ip during the solve, so you must pass capsolver your proxy credentials.
if you’re not sure, start proxyless. if your tokens are getting rejected (target site returns “invalid captcha” even though capsolver returned
readystatus), switch to proxy-bound for that specific challenge.import requests # proxy-bound recaptcha v2 task = { "type": "ReCaptchaV2Task", "websiteURL": "https://example.com", "websiteKey": "6Le-...", "proxyType": "http", "proxyAddress": "residential-proxy.example.com", "proxyPort": 8080, "proxyLogin": "user", "proxyPassword": "pass", "userAgent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36", } result = requests.post("https://api.capsolver.com/createTask", json={ "clientKey": "your-capsolver-key", "task": task, }).json()refund policy details
capsolver’s refund flow is simpler than most. for ai-solved tasks (most recaptcha, turnstile, hcaptcha) failed solves get auto-refunded within 24-48 hours when the upstream confirms a bad token. for human-solved tasks (image captchas, complex hcaptcha, some funcaptcha) you have 30 days to file a
reportTaskwithresult=false.requests.post("https://api.capsolver.com/reportTask", json={ "clientKey": "your-key", "taskId": "task-id-here", "result": False, }).json()practical advice: integrate the reportTask call into your scraper. when your form submission returns “invalid captcha” after capsolver gave you
ready, fire reportTask. you get the credit back without touching support. across a year of scraping, this saves a measurable amount.comparing total cost vs alternatives
service recaptcha v2 / 1k turnstile / 1k datadome / 1k capsolver $0.80 $0.80 $1.40 2captcha $2.99 $1.45 $2.99 anti-captcha $2.00 $1.30 $1.50 nopecha $1.00 $0.99 $1.50 deathbycaptcha $1.39 n/a n/a capsolver is consistently the cheapest. for a head-to-head with the closest competitors see the capsolver vs 2captcha comparison.
hidden costs to plan for
a few line items that don’t appear on the pricing page but show up in your real bill.
failed-token retries. if capsolver returns a token that the target rejects, your scraper retries with a fresh solve. budget 5-10% on top of the headline rate for retries even after refunds.
proxy costs for proxy-bound tasks. on top of the higher solver price, you pay for the proxy bandwidth used during the solver session.
enterprise upsell. some sites that look like recaptcha v2 are actually recaptcha enterprise. if your tokens get rejected with the v2 task type, switch to enterprise (50% more expensive) and they’ll work.
chrome extension solving. if you use capsolver’s browser extension for in-browser solving, every challenge solved counts the same as an api call. fine for testing, expensive at scale.
faq
how do i pay for capsolver?
prepaid wallet. top up with credit card, alipay, or major crypto (usdt, btc, eth, ltc, doge). minimum $5 top-up.does capsolver charge for failed solves?
not for ai-solved task types. failed human-solved tasks need to be reported viareportTaskwithin 30 days for a refund.what’s the cheapest captcha solver in 2026?
capsolver leads on most types. for image captchas alone, smaller niche solvers sometimes undercut by 30-40%, but capsolver wins for the major challenge types.is there a free trial for capsolver?
yes. free credits ($0.10-$1 depending on promotion) when you create a new account. enough to test 100-1000 solves before committing.can i set spend limits?
not directly via the dashboard. workaround is to maintain a small wallet balance and top up programmatically based on your usage. enterprise contracts include hard daily caps.what’s the typical solve time for capsolver?
recaptcha v2: 8-15 seconds. turnstile: 5-10. hcaptcha: 10-20. datadome: 12-25. ai-handled types are consistently fastest.conclusion
capsolver in 2026 is the price leader for captcha-solving across nearly every common challenge type. $0.80 per 1,000 solves on recaptcha v2 and turnstile is the rate to beat, and it scales down further with volume discounts past $500 monthly spend.
build the wallet alert. integrate the reportTask refund flow. start proxyless and only flip to proxy-bound when you need it. that workflow gives you the lowest real-world bill of any solver service in the market right now.
if you’re shopping solvers, this is your default. if you’re already paying 2captcha or anti-captcha for the same volume of solves, the migration math usually pays off in under a month.
-
How to Bypass Cloudflare Turnstile for Web Scraping (2026)
how to bypass cloudflare turnstile for web scraping (2026)
cloudflare turnstile is a captcha-replacement widget that issues a
cf-turnstile-responsetoken after fingerprinting the browser. you cannot solve it server-side with raw http. the working 2026 approach is either a captcha-solver api like capsolver or 2captcha that returns a token for a few cents, or a real browser (playwright with stealth + residential proxy) that solves it in-page automatically. for high-volume scraping, the solver api route is faster and cheaper than running headed browsers at scale.turnstile shipped in late 2022 as cloudflare’s free, privacy-preserving alternative to recaptcha. in 2026 it’s everywhere. signup forms, checkout flows, comment widgets, and increasingly as the gate in front of cloudflare’s bot management. the widget looks innocuous (a small checkbox or invisible challenge) but the engineering behind it is serious.
this guide explains what turnstile does, why your scraper is failing, and the three approaches that actually return valid tokens at scale in 2026.
what cloudflare turnstile does
turnstile sits as a
<div class="cf-turnstile" data-sitekey="...">on the target page. when the page loads, it pulls in the turnstile script fromchallenges.cloudflare.com/turnstile/v0/api.js. that script:- checks for “private access tokens” issued by apple, google, etc. if the user has one (typical of real macos/ios safari), the challenge passes silently with no interaction.
- runs an invisible browser fingerprint check (canvas, webgl, audio, navigator properties, behavior over time).
- if those fail, presents a managed challenge. visible widget, sometimes with a checkbox or interactive proof-of-work.
- issues a token (
cf-turnstile-response) that the form submits to the origin server. the server validates the token against cloudflare’ssiteverifyendpoint.
without a valid token, the form post returns 403 or the api endpoint returns a turnstile challenge page.
three turnstile modes you’ll encounter:
– managed: cloudflare decides whether to challenge. usually invisible.
– non-interactive: always invisible, lower friction.
– invisible: completely hidden from the user. challenge runs in the background.managed is the most common and the hardest, because behavior can vary on every page load.
why your scraper fails
if you’re seeing one of these symptoms, it’s turnstile:
– form submission returns a 403 withcf-mitigated: challengeheader
– json api returns html that containsJust a moment...orcf-turnstile
– requests with valid cookies still hit a challenge page on next call
– your selenium script that worked last week is now failing on the same formthree reasons standard scrapers fail.
first, the
cf-turnstile-responsetoken is required. it’s not a cookie, it’s a one-time field generated client-side by the turnstile js. you can’t fake it. you can’t reuse one. each form submit needs a fresh token tied to the page session.second, the turnstile script is heavily fingerprinted. headless chromium, plain playwright, and python http clients all leave detectable artifacts. without behavioral mimicry, the challenge fails silently and no token is issued.
third, ip reputation matters. requests from datacenter ips get a harder challenge. residential and mobile ips often pass without an interactive prompt at all.
the three approaches that work
ranked by reliability and cost.
approach 1: captcha solver api
the cheapest and fastest path for scale. you send the site key, target url, and user agent to a solver api. the solver returns a valid token. you submit the token with your form post. you never run a browser yourself.
capsolver is currently the price leader for turnstile. 2captcha and anti-captcha both work too. for the broader landscape see the best captcha solving services guide.
import requests import time CAPSOLVER_KEY = "your-capsolver-key" TARGET_URL = "https://example.com/login" SITE_KEY = "0x4AAAAAAA..." # extracted from the page's data-sitekey attribute def solve_turnstile(site_key, page_url): create = requests.post("https://api.capsolver.com/createTask", json={ "clientKey": CAPSOLVER_KEY, "task": { "type": "AntiTurnstileTaskProxyLess", "websiteURL": page_url, "websiteKey": site_key, }, }).json() task_id = create["taskId"] for _ in range(60): time.sleep(2) result = requests.post("https://api.capsolver.com/getTaskResult", json={ "clientKey": CAPSOLVER_KEY, "taskId": task_id, }).json() if result["status"] == "ready": return result["solution"]["token"] raise TimeoutError("turnstile solver timed out") token = solve_turnstile(SITE_KEY, TARGET_URL) print(f"got token: {token[:40]}...") # now submit the form with the token r = requests.post( "https://example.com/login", data={ "email": "user@example.com", "password": "test123", "cf-turnstile-response": token, }, ) print(r.status_code)cost: $0.0008-0.001 per solve at capsolver pricing in 2026. that’s roughly $1 per 1000 turnstile bypasses. faster than a headed browser, no browser footprint to maintain.
2captcha equivalent:
import requests import time API_KEY = "your-2captcha-key" def solve_turnstile_2c(site_key, page_url): r = requests.post("http://2captcha.com/in.php", data={ "key": API_KEY, "method": "turnstile", "sitekey": site_key, "pageurl": page_url, "json": 1, }).json() captcha_id = r["request"] for _ in range(60): time.sleep(3) check = requests.get( f"http://2captcha.com/res.php?key={API_KEY}&action=get&id={captcha_id}&json=1" ).json() if check["status"] == 1: return check["request"] raise TimeoutError()2captcha pricing in 2026 sits around $1.45 per 1000 turnstile solves. capsolver is consistently a few percent cheaper. at low volume the difference is rounding error.
approach 2: real browser with stealth
if you need to bypass turnstile during a fuller scraping flow (filling fields, clicking through a multi-step form, scraping content after the challenge), running a real browser is more natural than splicing in solver tokens.
import asyncio from playwright.async_api import async_playwright PROXY = { "server": "http://residential.example.com:8080", "username": "user", "password": "pass", } STEALTH = """ Object.defineProperty(navigator, 'webdriver', {get: () => undefined}); window.chrome = { runtime: {} }; Object.defineProperty(navigator, 'plugins', {get: () => [1,2,3,4,5]}); Object.defineProperty(navigator, 'languages', {get: () => ['en-US', 'en']}); """ async def main(): async with async_playwright() as p: browser = await p.chromium.launch( headless=False, proxy=PROXY, args=["--disable-blink-features=AutomationControlled"], ) ctx = await browser.new_context( viewport={"width": 1366, "height": 768}, user_agent="Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/131.0.0.0 Safari/537.36", locale="en-US", ) await ctx.add_init_script(STEALTH) page = await ctx.new_page() await page.goto("https://example.com/protected-page", wait_until="networkidle") # wait for turnstile to either pass silently or render the widget try: await page.wait_for_function( "() => document.querySelector('input[name=cf-turnstile-response]') && document.querySelector('input[name=cf-turnstile-response]').value !== ''", timeout=20000, ) token = await page.evaluate( "() => document.querySelector('input[name=cf-turnstile-response]').value" ) print(f"turnstile passed automatically. token: {token[:40]}...") except Exception: print("turnstile did not auto-pass, may need a click or solver fallback") # do whatever you came for html = await page.content() await browser.close() print(html[:2000]) asyncio.run(main())key choices:
–headless=False(or use xvfb for server). headless mode fails turnstile much more often than headed.
– residential or mobile proxies. datacenter ips push turnstile into harder challenges.
– the stealth init script patches obvious bot-tells.
– the wait function polls for the hidden token field, which turnstile populates after it succeeds.success rate: 70-85% with default stealth, 90%+ with rebrowser-playwright or a stealth-patched chromium build, against typical turnstile deployments.
approach 3: managed scraping api
if you don’t want to manage either solvers or browsers, services like scrapfly, brightdata web unlocker, and zenrows all handle turnstile transparently. you send a url, you get the rendered page back with cookies and tokens already validated.
from scrapfly import ScrapflyClient, ScrapeConfig client = ScrapflyClient(key="your-scrapfly-key") result = client.scrape(ScrapeConfig( url="https://example.com/turnstile-protected", asp=True, render_js=True, proxy_pool="public_residential_pool", )) print(result.content)cost: typically $1-3 per 1000 successful requests on cloudflare-protected urls. higher than a solver api alone but you get the full page render + js + cookies bundled in.
extracting the site key
the site key is the public identifier turnstile uses to know which widget config applies. you need it for solver api calls.
it’s hardcoded in the page html. open dev tools, search for
data-sitekey=orcf-turnstile. the value you want looks like0x4AAAAAAA.... it’s safe to hardcode in your scraper because it’s public.programmatic extraction:
import re import requests html = requests.get("https://example.com/login").text match = re.search(r'data-sitekey="([^"]+)"', html) if match: site_key = match.group(1) print(site_key)if the page loads turnstile dynamically via js, you’ll need playwright to find the rendered widget instead.
proxy choice for turnstile
proxy type turnstile pass rate typical cost per gb datacenter 30-50% $0.50-2 shared residential 60-75% $4-8 premium residential (bright data, oxylabs) 80-90% $6-15 mobile (4g/5g) 90-95% $10-30 mobile proxies are the highest-pass-rate option but the most expensive per gb. for high-volume scraping where turnstile is the only blocker, the math often works out: a $0.001 solver call beats a $0.05 mobile-proxy page load. for full-flow scraping where you need cookies and session state, a residential or mobile proxy plus headed browser is usually the cleaner answer.
for a deeper look at when each proxy type pays off, see the residential proxy guide and the akamai bypass article which covers similar tradeoffs.
hybrid approach: best of both
production scrapers usually combine the two cheaper options. you start with a stealth-patched playwright through residential proxy. if turnstile auto-passes, you keep going. if it doesn’t, you call the solver api as a fallback and inject the token.
import asyncio import requests import time from playwright.async_api import async_playwright CAPSOLVER_KEY = "your-key" async def get_token_from_page(page): try: await page.wait_for_function( "() => document.querySelector('input[name=cf-turnstile-response]')?.value", timeout=12000, ) return await page.evaluate( "() => document.querySelector('input[name=cf-turnstile-response]').value" ) except Exception: return None def solve_via_api(site_key, page_url): create = requests.post("https://api.capsolver.com/createTask", json={ "clientKey": CAPSOLVER_KEY, "task": {"type": "AntiTurnstileTaskProxyLess", "websiteURL": page_url, "websiteKey": site_key}, }).json() tid = create["taskId"] for _ in range(60): time.sleep(2) r = requests.post("https://api.capsolver.com/getTaskResult", json={ "clientKey": CAPSOLVER_KEY, "taskId": tid, }).json() if r["status"] == "ready": return r["solution"]["token"] return None async def main(): async with async_playwright() as p: browser = await p.chromium.launch(headless=False) page = await browser.new_page() await page.goto("https://example.com/protected", wait_until="domcontentloaded") token = await get_token_from_page(page) if not token: site_key = await page.evaluate( "() => document.querySelector('.cf-turnstile').dataset.sitekey" ) token = solve_via_api(site_key, page.url) await page.evaluate(f""" document.querySelector('input[name=cf-turnstile-response]').value = '{token}'; """) print(f"have token: {token[:40]}") await browser.close() asyncio.run(main())this gives you the speed of in-page autopass when it works and the reliability of solver fallback when it doesn’t. about $0.0002-0.001 per page on average, which beats either pure approach.
what doesn’t work in 2026
a few approaches that show up in older blog posts but are dead now:
- selenium with default chromedriver: navigator.webdriver detected instantly. fails 95%+ of turnstile challenges.
requestsplus copied cookies: turnstile tokens are single-use. cookies don’t carry between scrape sessions reliably.- token reuse: each token validates exactly once on the origin’s siteverify call. reuse returns 401.
- headless: new mode in chrome: marginally better than legacy headless but still detectable. not enough on its own.
- vpn-only setups (no proxy): most consumer vpns are flagged. residential proxies are different beasts.
faq
is bypassing cloudflare turnstile illegal?
not in itself. scraping public data is legal in most jurisdictions. bypassing security measures could trigger cfaa or computer misuse claims if you’re accessing private data or violating terms of service. the web scraping legal guide covers the case law.how much does it cost to solve turnstile?
solver apis charge $0.0008-0.0015 per token in 2026. for 1000 solves, expect $0.80-1.50 at capsolver, $1.40-1.50 at 2captcha.can i bypass turnstile with python requests alone?
no. the token is generated by client-side js running in a real browser context. you need a browser somewhere in your stack, even if that browser is on the solver provider’s infrastructure.why does my playwright script work locally but fail on a server?
servers usually run headless. headless chromium is more detectable. run with xvfb to fake a display:xvfb-run -a python script.py.does undetected-chromedriver bypass turnstile?
sometimes, with residential proxy. success rate around 60-70% on managed-mode turnstile. rebrowser-playwright is more reliable in 2026.what’s the difference between turnstile and recaptcha v3?
both are invisible-by-default scoring systems. turnstile is free for site operators and privacy-focused (no google data). recaptcha v3 is free for low volume but tied to google. solver pricing is similar, in the $0.001-0.002 range per token.conclusion
cloudflare turnstile is solved territory in 2026. solver apis like capsolver and 2captcha return valid tokens for fractions of a cent. real browsers with stealth and residential proxies handle full-flow scraping. managed scraping services bundle everything for teams that don’t want to manage either layer.
start with the cheapest approach that fits your use case. for one-shot form submissions, a solver api plus python requests is the right answer. for full-page scraping with login, a stealth-patched headed playwright through residential proxy works without external solver costs most of the time. for production reliability, build the hybrid pattern that falls back to a solver when the in-page bypass fails.
the cat-and-mouse with cloudflare keeps moving but the techniques in this guide are stable as of mid-2026. revisit your stack every six months.