Your cart is currently empty!
Retiring a scraper nobody uses
Fourteen scheduled jobs on one pipeline. I could name a human or a downstream consumer for nine of them. The other five had been running on a timer for somewhere between eight months and three years, and the honest answer to “who reads this” was that I had no idea.
That ratio is normal. Any pipeline alive longer than a year collects jobs whose reason for existing left the company with the person who asked for them, and the job carries on because a cron entry has no opinion about relevance.
The bill you can see and the bill you cannot
The visible cost is small enough to ignore, which is exactly why it gets ignored. One of those five fetched around 9,000 pages every six hours. On mobile addresses that is a real slice of a line I pay for monthly, plus the parsed table, plus the raw responses behind it, plus a nightly backup of both. Add it up and it is a rounding error against payroll.
The invisible cost is the one that hurts.
An old job breaks occasionally. A proxy goes bad, a template shifts, a run times out. That fires an alert, and somebody opens the alert, recognises the job name, and closes it. Second time they close it faster. By the third month they close it without reading it, and that reflex does not stay politely attached to the one job that trained it. You have taught a person that alerts from this pipeline are noise.
Then there is incident time. When the proxy pool degrades or a queue backs up, the dead job is right there in the graphs near the top of the request count, and somebody has to rule it out before reaching the actual fault. I have spent 40 minutes of a two hour incident eliminating jobs that turned out to have zero readers. That is the price: skilled attention at the worst possible moment, several times a year.
And one cost lands entirely on someone else. Those runs sent roughly three quarters of a million requests to a company that got no order, no signup, and no ad impression out of any of them. Public pages, robots file honoured, a pace slow enough that it would not have registered on their side. All true, and it was still three quarters of a million requests in service of a table with four reads in it.
Why the decision never gets made
Nobody turns these off because nobody can prove they are unused, and the cost of guessing wrong feels lopsided.
Delete it and be wrong, and you have quietly broken something for a colleague who will discover it at month end. Keep it and be wrong, and you pay a few dollars a month that nobody measures. Given that asymmetry, “keep it” wins every review, in every quarter, indefinitely.
The framing is the problem. You are being asked for a judgement call while the evidence is missing, and the correct response is to go and collect the evidence rather than argue about the odds.
A green run is a write side metric
Every scraper I have built reports on itself: runs completed, rows written, error rate, duration, and a chart that sits above 99%. All of that describes what happened on the write path. It confirms the code executed and the rows landed where they were pointed.
It says nothing about whether anybody wanted them.
That is the confusion keeping zombie jobs alive. From the inside, a healthy job and a useful job look identical. A scraper cannot distinguish between feeding a live pricing model and filling a table nobody has opened since last March, because both of those runs are green.
Watching for silent failures in a job that matters is a separate discipline with its own signals, and I have written about it separately. This is the prior question. Before you invest in monitoring a job properly, establish that the job has a reader.
Count the reads
Reads are the only real evidence of use. Everything else is inference with confidence added.
So measure them per table, per file, per endpoint, with an identity attached wherever the platform will give you one.
In Postgres, the counting is already happening. pg_stat_user_tables keeps a per table tally of sequential and index scans that increments on every read. It resets when the server restarts, so snapshot it on a daily cron and store the delta. Two weeks of deltas gives you something to argue from.
On object storage, switch on bucket access logging and retain 90 days. Same signal, cheaper.
Behind an API, log the caller identity on every request. A count tells you the table gets read. An identity tells you who to go and talk to.
Feeding a dashboard, most reporting tools already record a “last viewed” timestamp and viewer per chart. That field is usually sitting unused in a metadata table.
Exclude your own plumbing first
This trap will waste a month if you let it. Your own infrastructure reads the table constantly. The nightly backup reads it. Replication reads it. A health check running SELECT 1 against it every minute reads it, and that alone will report 1,500 reads a day from a table no human has touched.
Give those a dedicated database user or API key and filter that identity out of the count. Twenty minutes of setup, and without it the entire exercise returns noise.
What should remain is a person running a query, or another job in your own pipeline consuming the output. Those two count as use. Everything else is your plumbing talking to itself.
If you cannot instrument reads at all, the crude version still works: switch it off and see who shouts. I have used it. It only stops being reckless once you have announced it, because otherwise you are breaking things and calling it research.
Pause before you delete
Even holding the numbers, people hesitate, and the hesitation is rational. Deletion is the one operation with no undo.
So do not open with deletion. Four steps, and the first is free.
Stop the schedule. Comment out the cron line, pause the DAG, disable the trigger. Code stays. Data stays. Nothing is destroyed and the reversal is a single line.
That step alone captures nearly all the benefit. Requests stop, alerts stop, the job disappears from the incident graphs, and no permanent decision has been made.
Announce it. Wait a defined window. Then delete the schedule and the job definition, and stop there.
Word the announcement as a decision
Put it where your team already argues about deploys, not on a wiki page nobody has opened since onboarding.
And phrase it as a decision with a date rather than a question. Ask “does anyone use this?” and you get silence, silence reads as risk, and the job runs for another year. Say instead: this job is paused, the table stays, and unless something comes up it gets deleted on the 15th. Now silence is an answer, because you stated the default and gave people a cheap way to override it.
That single change of wording has unstuck more of these for me than anything technical.
What survives the deletion
Keep the data. Keep enough of the code to explain how the data was produced.
Somebody will ask, two years from now, what a column in an old report meant. Whether the price included tax. Why the figures jump in one month. The answer to all of that lives inside the parser, and if the parser is gone the answer went with it.
The case for versioning a scraper alongside its output and holding raw responses is made elsewhere on this site, and I am not going to relitigate it here. This is the same argument arriving from the other end. Code is the documentation for data, so the code outlives the schedule that ran it. A repository with a dead job in it costs a few megabytes.
Where a short window lies to you
Here is the one I got wrong, and it is worth more than the rest of this.
I instrumented reads on a job for 30 days, saw nothing outside my own queries, paused it and deleted it a month later. Clean process, exactly the steps above. In January somebody went looking for the year end comparison and found the table stopped in October.
Thirty days of read data proves nothing about anything quarterly, and nothing whatsoever about a report that runs annually. Every business I have worked with has at least one of those.
So the window has to exceed your longest reporting cycle, which for most people means a full year, and nobody will wait a year before turning something off. The way out is the pause. Pause for the year rather than run for the year. Pausing costs nothing, the history stays intact, and if December brings a complaint you restore a cron line and lose an afternoon.
The honest limit: a pause is reversible for the code and not for the calendar. Months spent paused have no rows in them and never will. On a time series that is already a partial deletion, and you owe people that detail in the announcement.
The position
A scraper you cannot prove is used costs more in attention than it costs in servers, and attention is the resource you are actually short of.
I would rather delete a job somebody quietly needed and rebuild it than carry twenty I cannot account for. Rebuilding a collector against a site you have already solved is two days. Twenty unexplained jobs is a permanent tax on every incident and every new person you hand the pipeline to. The rebuild is only cheap because you kept the data, which was always the expensive half, and which the procedure above is designed to protect. More on pipeline hygiene and the proxy infrastructure underneath it is at dataresearchtools.com.
Get new guides and videos first — join the Telegram channel.
Leave a Reply