Your cart is currently empty!
Handling pagination that changes shape
The problem shows up as a silent data gap, not an error
Most scraping failures are loud. A 403, a CAPTCHA wall, a timeout. Pagination failures are quieter. The scraper keeps running, keeps returning 200s, and keeps writing rows to your database. It just stops writing the rows that matter, because it fell off the end of a pagination scheme it assumed would hold steady. You find out three weeks later when someone asks why a category page that should have 40,000 listings only has 1,200.
This happens because pagination logic gets written once, against whatever the site is doing on the day you build the scraper, and then treated as a fixed contract. It isn’t one. Pagination is a UI and infrastructure decision made by product and backend teams for reasons that have nothing to do with you, and it changes more often than scrapers assume.
The shapes pagination actually takes
There are really four families, and most sites use more than one depending on where you are in the site.
Offset and page number. ?page=3 or ?offset=60&limit=20. This is the easiest to reason about and the easiest to break, because it assumes a stable, ordered list underneath. If the site re-sorts results between requests (common on anything with “sort by relevance” or live inventory), offset pagination silently skips or duplicates items even without any format change.
Cursor-based pagination. Instead of a number, the response includes an opaque token, something like next_cursor: "eyJpZCI6NDIw...", that you pass back on the next request. This is what most modern APIs use, including GraphQL’s Relay-style edges / node / pageInfo.endCursor pattern. It’s more resistant to re-sorting but it means your scraper can’t jump to page 40 without walking through 39 first, and it can’t be parallelized the same way offset pagination can.
Infinite scroll / “load more”. No URL parameter at all from the user’s point of view. The page fires an XHR or fetch call on scroll or button click, usually hitting an internal API that mirrors one of the two patterns above. The catch is that this internal API is often undocumented, versioned separately from the public one, and changed without any changelog.
Hybrid and inconsistent. The one that actually causes the most incidents. Category pages use offset, search results use cursor, and the mobile web view uses infinite scroll hitting a third endpoint entirely, all on the same domain. A scraper built against one page type and pointed at another silently mispaginates.
Why sites change shape mid-flight
It’s worth being clear-eyed about this: most pagination changes are not anti-scraping measures. They’re normal engineering. A team migrates a REST endpoint to GraphQL. A backend swaps offset pagination for cursor pagination because deep offsets get slow on large tables (this is a genuine, well-documented Postgres and MySQL performance problem, not a scraper deterrent). A frontend redesign replaces numbered pages with infinite scroll because it tests better on mobile. An A/B test rolls out to 10% of traffic and your scraper happens to land in the test bucket on one run and the control bucket on the next.
Some of it is deliberate friction. Sites that are aware of heavy automated traffic do sometimes make pagination less stable on purpose, rotating parameter names or adding required session state between page requests, specifically because static scrapers assume URL patterns hold. But you generally can’t tell which cause you’re looking at just from the symptom, and treating every change as hostile leads to fragile, paranoid code. Treat it as what it almost always is: the site changed and your assumption didn’t update with it.
Detect the shape instead of assuming it
The fix starts before you write a single line of extraction logic. Open the network tab, not the rendered HTML, and find the actual request that returns the list data. Look at three things:
- Does the request include a page number, an offset, or a cursor token in the request itself?
- Does the response include a total count, a next-page URL, a next-cursor field, or none of those?
- Is the endpoint versioned (
/api/v2/listings) or does it look like an internal, unversioned call that could change without notice?
If the response has an explicit has_next_page: false or an empty next_cursor, your stopping condition is right there in the data and you should use it instead of guessing. If it doesn’t, you’re dealing with a site that expects the frontend, not an API consumer, to decide when to stop, which is a signal that the format is more likely to shift under you later.
Building a scraper that survives a format change
The practical fix is to stop hardcoding a single pagination strategy and instead write an extraction step that checks its own assumptions on every run, not just at build time.
Concretely, that means:
- Parse the pagination signal from the response itself (cursor field, next-page link, total count) rather than incrementing a counter you set once in config.
- Add a schema check at the top of your parsing function. If the expected
next_cursorkey is missing or a new key likenextPageTokenappears in its place, throw a specific, loud error instead of falling through to a default that silently stops paginating. - Log the item count returned per page alongside the page identifier. A page that suddenly returns 0 items when every prior page returned 20 is a stopping condition, but it’s also worth flagging separately from “we hit the end,” because it could mean the endpoint started rejecting your requests in a way that still returns 200.
- Cross-check total item count against what the site’s own UI reports, when it reports one at all. A scraper that pulls 1,200 of a stated 40,000 listings and exits cleanly is the exact failure mode that goes unnoticed the longest.
None of this requires exotic tooling. It’s mostly about not trusting your own code’s silence. A pagination loop that can exit without an explicit, logged reason is the thing to eliminate.
Stopping conditions that don’t lie to you
A surprising number of pagination bugs come down to a stopping condition that’s technically correct and practically wrong. “Stop when the page returns fewer than the requested limit” is a reasonable rule until a site starts padding short pages with recommended or sponsored items to keep the count consistent, at which point your scraper never stops and you’re re-requesting the same tail page in a loop.
“Stop after N pages” is worse, because it has no relationship to the actual data and will quietly under-collect the moment the site’s result set grows.
The stopping conditions that hold up are the ones grounded in something the site itself asserts: an explicit has_more: false, a next_cursor that comes back null, or a total count you can check your running total against. When none of those exist, the more honest approach is to paginate until you see a repeat of an item you’ve already collected (by ID, not by position), which tells you the underlying list looped or reset, and to log that event rather than treat it as a normal end state.
What this means for pipeline design
If pagination logic lives buried inside a single scraper script, a format change means a silent partial run every time until someone notices the row count looks wrong. The more durable pattern is to treat pagination parsing as its own small, tested unit, separate from request handling and separate from data transformation, so a change in the site’s response shape throws a clear, specific failure in one place instead of a vague downstream data-quality problem three steps later. Combined with per-run item count logging and count checks against any total the site publishes, this turns “pagination changed and we didn’t know” from a recurring incident into a one-line alert.
Pagination that changes shape is not a special case. On any site actively maintained by a real engineering team, it’s the default state over a long enough timeline. Building for that from the start costs a little more up front and saves you from finding gaps in your data months after they opened.
If you want more of this kind of infrastructure-level breakdown of how scraping pipelines actually hold up in production, you can find the rest of our write-ups on the Data Research Tools home page.
Get new guides and videos first — join the Telegram channel.
Leave a Reply