What Web Scraping Actually Costs (2026)

The first time a client asked me what it would cost to scrape a site, I gave them a number off the top of my head, and it was wrong by a mile. I had priced the one thing everyone prices, the proxies, and forgotten almost everything else. By the time the job was running, the real bill was close to three times my guess, not because anything went wrong, but because a scraper costs money in five different places and I had counted one of them.

I run proxy infrastructure and production scrapers for a living, so I have paid every one of these bills myself, more than once. This is not a pricing page I copied. It is where the money actually goes when you run this at scale, and which parts of it people forget until the invoice lands. Here is the whole picture in one table before the detail.

Cost bucket What you pay for Usually billed Easy to forget?
Proxies and bandwidth The traffic that carries your requests Per gigabyte or per request No, everyone prices this
Compute The machines that run your code Per hour of CPU and memory Sometimes
Unblocking Solvers and scraping APIs when a site pushes back Per solve or per request Sometimes
Storage and pipeline Where the data lands, lives, and gets queried Per gigabyte, per month Often
Maintenance The human hours that keep it all alive Engineer time Almost always, and it is the biggest

Most people quote the first row and pretend the other four are free. That is exactly the mistake I made.

Bucket one: proxies and bandwidth

Proxies are the bucket everyone fixates on because it is the easiest to see a price for. Residential and mobile addresses are usually sold by the gigabyte, so every page you pull, and every image and script it drags along, adds to a running meter. Datacenter addresses are cheaper and often sold by the request, but they get blocked far more often, which pushes the cost somewhere less obvious. This bucket is real, but it is almost never the whole story, and treating it as the whole story is how you underquote by a factor of three.

Cost per successful page, not cost per gigabyte

The trap inside the proxy bucket is thinking in cost per gigabyte when you should think in cost per successful page. A cheap datacenter address looks like a bargain until half your requests come back blocked and you fetch each one again, paying twice or three times over for a single page you keep. A pricier mobile address that lands on the first try can be cheaper per good page than the bargain that fails constantly. Do not shop on the sticker rate. Shop on what it costs to end up with one clean page you can use, because that is the number your budget actually feels.

Bucket two: compute

Compute is where the browser question quietly decides your bill. A plain HTTP client that fetches raw HTML is almost free per page. It barely touches memory or processor time, so you can run many of them on a small cheap machine. A real browser is the opposite. It eats memory and processor time for every page it opens, so a collector built on browsers can cost you many times what the same job costs with a plain client. On a cloud bill that difference is not a rounding error, it is often the largest line you have.

The lever is simple, and it is the same one I preach everywhere. Reach for a real browser only on the pages that genuinely need one, where the data is painted on by JavaScript and is not in the raw HTML. For everything else, usually most things, a plain client that grabs the fully formed HTML is faster and dramatically cheaper. Reaching for a browser by default, on pages that never needed it, is the single most common way I watch a compute bill balloon for no reason.

Bucket three: unblocking, and an honest word on detection

Unblocking is the bucket people either overspend on or refuse to think about. CAPTCHA solving services charge you per solve. Scraping APIs charge per request, but they bundle the proxies, the browser, and the retries into one price so you manage none of it yourself. That convenience is genuinely worth paying for on some jobs. The point here is the cost shape: you are paying a premium for someone else to absorb the hard part, and that premium is easy to lose track of when it is buried inside a per request rate.

I have to be honest about this bucket, because it is where people get the wrong idea. None of these services make you undetectable, and none of them buy you the right to take something you were not allowed to take. They help you fetch public pages more reliably, that is all. You still move at a gentle pace, respect the robots file and the site’s terms, and take public data only. The deeper you go into fighting a detection team that ships updates faster than you do, the more this bucket costs and the less durable your setup becomes. The cheapest long run path is almost always the one that is not adversarial in the first place.

Bucket four: storage and the pipeline

Storage is everything that happens to the data after you pull it. You keep the raw pages, the cleaned and parsed results, probably a database or a warehouse to query them, and the machinery that moves data from one stage to the next. None of it is huge on day one, but it grows every day the job runs, and it is a recurring cost, not a one time purchase. A scrape you run nightly quietly multiplies its storage and processing by the number of nights, and that slow accumulation is a part people forget entirely when they budget.

There is a small nasty surprise hiding in this bucket too. Moving data around costs money in ways that are easy to miss. Many cloud providers let you bring data in for free but charge you to take it back out, so the day you export your results or hand them to a client, a bill you never saw coming can appear. It is rarely the biggest number, but it is exactly the kind that ruins a tidy estimate, so leave a little room for it.

Bucket five: maintenance, the one that dwarfs the rest

Maintenance is the bucket that dwarfs the other four on any job that runs for more than a few weeks. Websites change their layout, and when they do your selectors point at nothing and start collecting blanks. Addresses get blocked and pools need topping up. A site adds a new defence and your collector needs adjusting. Every one of those is a human sitting down and fixing something, and human hours are the most expensive thing in this whole picture by a wide margin. A scraper is not a thing you build once, it is a subscription of attention, and the attention is what you are really paying for.

This is also why the cheap tool is so often the expensive one. A fragile collector that saves a little on proxies but breaks every other week, and pulls an engineer off other work each time, costs far more in the end than a slightly pricier setup that just runs quietly. When you count the cost of a scraper, count the hours, not only the dollars on the cloud invoice, because on a long running job the hours are the invoice.

Build cost versus run cost

The single framing that fixes most bad estimates is separating the build cost from the run cost. Building a scraper is a one time job. You write it, test it, get it working, and that is a fixed number of hours you pay once. Running it is a completely different thing, because the proxies, compute, storage, and maintenance all land again on every execution, night after night. People quote the build and forget the run, and the run is where nearly all the money lives over the life of the job. Always price them as two separate things.

Which bucket dominates depends on scale

Which bucket dominates depends entirely on the size of your job, and knowing your regime tells you what to optimise. On a small job, a few thousand pages here and there, the infrastructure is basically free and your own time is the whole cost, so build it fast and do not over engineer it. On a huge job, millions of pages a day, the proxies and compute become the giants and shaving a little off cost per page saves real money. Optimising the wrong bucket for your scale is wasted effort, so work out which regime you are in before you tune anything.

How I estimate a job

When someone asks me to estimate a real job, I walk the five buckets one at a time. I guess how many pages a day, and whether they need a browser or a plain client, which sets the compute. I estimate a cost per successful page for the proxies, allowing for retries, not the sticker rate. I add whatever solving or API help the site will force on me, and storage that grows over time. And then, most important, I add a standing allowance for maintenance, because I know the site will change and something will break. It is rough, but it is honest, and it is never off by a factor of three.

The cheapest scrape is the one you never run

The best cost lever of all is the one people skip: simply not fetching pages you do not need. Cache what you already have so you never pay for it twice. Fetch only what has changed since last time instead of pulling the whole site again every night. Hit the quiet background request that hands you clean data directly, so you skip the heavy browser entirely. Every page you avoid is pure saving across all five buckets at once, no proxy, no compute, no storage, no maintenance. The cheapest scrape in the world is the one you were smart enough not to run.

I run this infrastructure in production myself, so none of this is theory, it is the bill I actually pay. If you want the full written breakdown of all five buckets, the way I estimate a job before I quote it, and the proxy pool I route my own scrapers through, it is all here, with no magic numbers and no promise that your job will cost exactly what mine does, because that is not a thing anyone can honestly tell you without seeing your pages.

Get new guides and videos first — join the Telegram channel.

Comments

Leave a Reply

Your email address will not be published. Required fields are marked *