Your cart is currently empty!
Personalised prices and why two scrapers disagree
What personalised pricing actually is
You run the same scraper against the same product page twice, ten minutes apart, and get two different prices. Nothing crashed. No error in the logs. Just two numbers that don’t match. If you’ve spent any time monitoring competitor pricing at scale, you’ve hit this. The instinct is to assume a bug in your own code. Often the real answer is that the site never had one price to begin with.
Personalised pricing means the price (or the discount, the shipping estimate, the “only 2 left” badge) is computed per visitor rather than per product. The product record in the backend has one canonical price, but what gets rendered depends on signals attached to the request: who’s asking, from where, on what device, with what history. Airlines and hotel booking sites have done this for years. Retail has caught up, partly because it’s cheap to do once you already have an experimentation platform in place for UI testing.
For anyone running scrapers for price intelligence, MAP monitoring, or competitive research, this matters because it breaks the basic assumption that “the price” is a stable fact you can pull once and trust.
The mechanics behind price personalisation
Personalisation engines don’t usually hardcode “charge this user more.” They work off a bucketing function: some identifier (a cookie ID, a device fingerprint hash, sometimes IP-derived geo) gets fed into a hash, and the hash output determines which variant of the page you see. This is the same infrastructure used for ordinary A/B testing on layout or copy, just pointed at price or discount logic instead.
Signals that commonly feed the bucketing decision:
- Session and cookie history (returning visitor vs new, cart abandonment)
- Geolocation resolved from IP address
- Device and browser fingerprint (screen size, user agent, sometimes canvas/WebGL signals)
- Account state if logged in (loyalty tier, past purchase volume)
- Referrer (arrived from a coupon site vs direct navigation)
- Currently running experiments the traffic gets randomly assigned into
None of these are exotic. They’re the standard inputs any web analytics stack already collects. The personalisation layer just reuses them to decide which price variant to serve.
Why your two scrapers see different numbers
This is the part that trips people up: two scraper runs against the “same” page are rarely identical requests. Even with the same target URL, they can differ on almost every signal the bucketing function reads.
If your scraping infrastructure rotates proxies between runs, you’re changing the resolved geolocation and often the ASN, which can move you into a different regional pricing zone or a different experiment cohort entirely. If cookies aren’t persisted between runs, every scrape looks like a brand new anonymous visitor, which itself is a bucketing signal (new visitors sometimes get a welcome discount that returning visitors don’t). If your headless browser configuration varies run to run (different viewport, different user agent string picked from a rotation list), the fingerprint hash changes, and so does the bucket.
None of this requires the site to be doing anything adversarial toward scrapers specifically. It’s the same bucketing logic that would show two real human shoppers on two real devices two different prices for the same reason. Your scraper just happens to be triggering the same mechanism from run to run because its own identity isn’t held constant.
A/B testing buckets vs genuine personalisation
Worth separating two things that get conflated. Genuine personalisation ties the price to something durable about you: your account, your loyalty tier, your purchase history. A/B testing ties it to a random assignment that’s supposed to be stable for a given identifier but arbitrary in what it means. If you’re bucket A, you get the control price; if you’re bucket B, you get the 10% off variant. Neither bucket is “the real price.” Both are live in production at the same time, on different slices of traffic, so the business can measure which one converts better.
From a scraper’s point of view these look identical: two runs, two prices, no visible reason. But the practical implication differs. Genuine personalisation might mean a real customer profile would consistently see one price, so a competitor intelligence report built on scraped data should say “we observed a personalised range of X to Y” rather than pretending there’s one ground truth number. An A/B test in progress means the price is temporary and will resolve to one variant once the test ends, and sampling it once tells you almost nothing about where it’ll land.
The role of caching and edge nodes
A layer that gets ignored: CDN and edge caching. Sites using Fastly, Cloudflare, or similar often cache rendered pages at edge locations close to the requester. If pricing is baked into the cached HTML rather than fetched client-side via an API call, then which edge node answered your request matters. A proxy exit in Frankfurt and one in Singapore can hit different cache tiers with different TTLs and different regional price configurations, independent of any personalisation logic at all.
This is a separate failure mode from bucketing, and it’s worth ruling out first because it’s the easiest to diagnose: check response headers for cache status (age, x-cache, cf-cache-status and similar), and check whether the price appears in the initial HTML payload or gets fetched by a subsequent XHR/fetch call to a pricing API. If it’s the latter, that API call is a more reliable point to instrument than the rendered page, since it typically carries fewer caching layers.
Building a scraper that gives you a stable baseline
If the goal is to track a competitor’s price over time, the fix isn’t to fight personalisation, it’s to hold your own request identity constant so you’re comparing like with like. In practice that means:
- Pinning a proxy pool by region deliberately, not rotating randomly between requests when consistency matters more than volume
- Persisting cookies and session state across scheduled runs instead of starting fresh each time
- Keeping browser fingerprint parameters (user agent, viewport, timezone) fixed for a given monitoring identity, rather than randomising per request
- Logging the response headers and request signals alongside the scraped price, so a later price change can be traced back to “new proxy IP” or “cache expired” rather than treated as a mystery
If the goal is closer to market research, understanding the distribution of prices a real customer base actually sees, then the opposite approach applies: deliberately vary geo and session state across multiple identities and treat each scrape as one sample from a distribution, not a single fact. That only works with a large enough sample to say anything statistically meaningful. A handful of runs from a handful of IPs isn’t a distribution, it’s noise with a story attached.
What you can conclude and what you can’t
A single scrape showing a different price than yesterday’s scrape is not evidence of an active A/B test, and it’s not evidence of nothing either. It’s a prompt to check what changed on your side of the request first: proxy exit, session state, fingerprint, cache headers. Only once those are controlled for does a repeated, reproducible price difference across otherwise-identical requests start to look like genuine personalisation worth reporting on.
It’s also worth being honest in any report or dashboard about what’s a measured fact and what’s an inference. “This price varies by region, confirmed across N runs from M distinct pinned locations” is a claim you can stand behind. “This site personalises price by loyalty status” is a much stronger claim that needs an actual logged-in test to support it, not a guess from anonymous traffic.
The ethical and legal edges worth respecting
Price pages that are publicly visible without login are a different category from anything gated behind an account, and personalisation research should stay on the public side of that line. Scraping a site’s published, non-personal pricing to understand a market is a common and legitimate use of scraping infrastructure. Pulling data tied to a specific logged-in account, or a real individual’s personalised offers, isn’t the same thing, and it’s the kind of line worth being deliberate about before a project starts, not after. Whatever a site’s terms of service say about automated access is also worth reading before building a monitoring pipeline against it, since the answer varies a lot by site and isn’t something a scraping tool can decide for you.
If you’re building out price or A/B test monitoring and want to compare notes on proxy infrastructure, session handling, or pipeline design, that’s exactly the kind of thing we write about here.
Get new guides and videos first — join the Telegram channel.
Leave a Reply