Your cart is currently empty!
Deciding how often a page really needs rechecking
The default interval is almost always wrong
Most scraping setups start with one number. Every hour, every six hours, every day, applied to every target in the list regardless of what that page actually is. It’s the easiest thing to configure and the first thing that gets abandoned once the pipeline is running in production, because nobody revisits it until something breaks.
The problem is that “how often does this page change” is not one answer. A product listing page, a company’s about page, and a job board posting all update on completely different clocks, and treating them the same wastes either compute or freshness, usually both at once.
What you’re actually paying for on every check
Every recheck costs something even when nothing changed. A proxy request, a render if the page needs JavaScript, storage for the response, and a diff pass to confirm there was no update. None of that is free, and at scale it adds up faster than people expect.
Run the math on a target list of a few thousand URLs checked hourly. That’s tens of thousands of requests a day before you’ve extracted a single new fact. If 90% of those checks return an unchanged page, which is common for slow-moving content, you’re burning proxy bandwidth and render time to confirm nothing happened. It’s the majority of your operating cost on a target that didn’t need it.
On the other side, checking too infrequently means you miss the window entirely. A flash sale that runs for six hours, a stock alert that clears in minutes, a job posting pulled after the role fills. If your interval is longer than the thing you’re tracking is even live, the frequency number is decorative. You paid for infrastructure that structurally cannot catch the event.
Group targets by how they actually behave, not by convenience
The fix isn’t a smarter algorithm, it’s admitting that your target list is not one population. Split it by observed change behavior and give each group its own interval.
Near-static pages. Company about pages, terms of service, static documentation, archived listings. These change on the order of weeks or months. Checking them hourly produces months of identical diffs. A weekly or even monthly check catches essentially everything that matters, and the savings on request volume are large because this is usually the biggest bucket by page count.
Slow-moving content. Product catalog pages without live inventory counters, blog indexes, category pages. These shift daily or every few days as new items get added or descriptions get edited. A check every 6 to 24 hours is usually enough, and you can tighten that if you notice a specific source publishing on a predictable schedule.
Fast-moving content. Pricing pages, inventory counts, job boards, auction listings, anything with a countdown or a limited quantity. These need checks measured in minutes, not hours, and even then you’re accepting that you’ll sometimes land between two states of a fast-changing value.
The grouping doesn’t have to be manual forever. Once you’ve run a target for a few weeks, you have real diff history. Use it.
Let observed change rate set the interval, not a guess
The most reliable signal you have is the page’s own history. If you’ve been storing diffs (and you should be, even a simple hash of the extracted fields is enough), you can calculate the actual interval between changes for each URL instead of assuming one.
A basic version of this: track the timestamp of the last N changes for a page, take the median gap between them, and set your check interval to some fraction of that, say a quarter to a half. A page that changes roughly every 3 days doesn’t need hourly checks, but it does need more than a weekly one if you want to catch the change within a day of it happening. A page that hasn’t changed in 60 days probably doesn’t need daily checks anymore even if it started in your “daily” bucket.
This is adaptive scheduling, and it doesn’t require anything exotic. A cron-based scheduler with a per-target interval column in your database gets you most of the way there. The part people skip is updating that column based on observed behavior instead of setting it once at ingestion and forgetting it.
Respect what the source is telling you
Some pages are explicit about their own update cadence, and it’s worth reading before you set an interval. RSS and Atom feeds sometimes expose a ttl or updatePeriod field. Sitemaps can include a lastmod timestamp per URL that, if the site maintains it honestly, tells you exactly when a page last changed without you having to poll it at all. robots.txt occasionally sets a Crawl-delay directive, which isn’t a change-frequency signal but is a direct statement from the site operator about how much load they’re willing to absorb from automated requests, and ignoring it is a fast way to get an IP range flagged.
Checking a sitemap’s lastmod values on a schedule is often cheaper than checking the pages themselves, because it’s one lightweight request that can tell you which of a thousand URLs actually need a full recheck this cycle. Not every site keeps this field accurate, some just stamp it with the current date on every deploy regardless of content changes, so treat it as a hint to prioritize, not a guarantee.
Where request timing intersects with detection
This isn’t just a cost question. Request timing is one of the signals that bot detection systems look at, alongside things like TLS fingerprint consistency, header ordering, and browser automation artifacts. A target hit at exactly the same second every hour, from the same IP or IP range, forever, produces a pattern that’s trivial to flag on the server side regardless of how good your headers look. Detection systems don’t need to catch a single request looking wrong, they just need a request pattern that a real user’s browsing behavior wouldn’t produce.
Adding jitter to your interval, checking every “roughly 60 minutes” with random variance of a few minutes in either direction, is a defensive move against exactly that kind of pattern-based fingerprinting, not a workaround for anything on the target site. It’s the same logic as varying your proxy pool: consistency in timing is itself a signal, and reducing it reduces one input into a detection system’s decision, nothing more. It doesn’t make a request undetectable and it isn’t a substitute for respecting a site’s published crawl policies or terms of service.
A rough starting framework
If you’re setting this up from scratch, here’s a reasonable place to start before you have your own diff history to lean on:
- Static/legal/about pages. weekly checks
- Catalog, blog, listing pages without live counters. every 6 to 12 hours
- Pricing, inventory, time-limited listings. every 15 to 60 minutes, with jitter
- Anything with a sitemap
lastmodfield. check the sitemap first, only recheck the page if the timestamp moved
Revisit these numbers after your first month of real diff data. The point of the exercise isn’t to land on the perfect interval on day one, it’s to stop treating every URL like it belongs to the same clock.
If you’re building out the orchestration side of this, from scheduling to proxy rotation to the fingerprinting considerations that come with running requests at scale, that’s the kind of infrastructure we write about and review at Data Research Tools.
Get new guides and videos first — join the Telegram channel.
Leave a Reply