I set myself a real-world test: collect every apartment listed for sale in one city, from a large property portal sitting behind a Cloudflare challenge, and deliver it the way a paying client would expect.
The site's own counter said 141 listings.
I delivered 109.
And I could prove that 109 was 100% of what the site actually publishes.
This post is about that gap, because it's the question every buyer of scraped data should ask, and almost nobody does: how do you know you got everything?
THE WALL WAS NOT THE HARD PART
The first request came back 403. So did the homepage. The site was running a JavaScript challenge that refuses anything that doesn't behave like a real visitor.
Getting through took a session that respected the site's own rules: its crawling rules, a human pace, no forbidden paths, and no personal data collected at any point. Once that was in place, the listings loaded.
That's where most scraping jobs stop. Rows in a file, job done.
It's also exactly where the expensive mistakes start.
141 VS 109: THE QUESTION THAT MATTERS
A gap of 32 listings. Two possible stories:
- My scraper missed 32 listings. The dataset is incomplete, and every average, every price per m² built on it is quietly wrong.
- The counter overstates what the site actually serves. The dataset is complete, and the counter is the thing that's off.
From the outside, both look identical. A file with 109 rows. The only way to know is to check.
So I didn't trust a single path. I collected the same city through three independent entry points, then compared them line by line.
All three converged on the same 109 listings. Not one listing appeared in one path and not the others.
The counter was counting things the site doesn't actually publish in its lists. 109 was the full set. Nobody has to take my word for it: the proof comes with the file.
THE SILENT TRAPS ALONG THE WAY
None of these threw an error. All of them would have shipped a wrong dataset:
- Number formats. The listings write 72,30 m², with a decimal comma. Read the wrong way, that becomes 7,230 m². One apartment the size of a shopping mall, and the average surface of the whole city is garbage.
- Infinite scroll that gives up early. The list loads more results as you scroll. Stop a little too soon and you silently collect 11 listings, or 29, instead of 109.
- Partner ads mixed into the results. Some cards look exactly like listings but come from other sites. Count them and your totals are inflated; drop them blindly and you lose real data.
- A city page that isn't only the city. Nearby towns get mixed in with distances attached. Fine if you know it, misleading if you don't.
Every one of these is invisible if you only check that the job finished.
WHAT A CLIENT ACTUALLY BUYS
Not 109 rows. Anyone can hand over rows.
A client buys certainty: a clean file, one row per real listing, every field normalised, and a short quality report showing coverage, missing fields per column, and how the total was verified.
That report is the difference between data you can build decisions on and data you hope is right.
TAKEAWAYS
- The wall is rarely the hardest part. Proving completeness is.
- A site's own counter is not ground truth. Verify against what it actually serves.
- Silent traps don't crash. They just make your numbers wrong.
- Ask your data supplier one question: how do you know nothing is missing?
I build and maintain scrapers for protected and fast-changing sites, and every delivery comes with proof of what was collected and what wasn't. If you're not sure your dataset is complete, or your scraper keeps getting blocked, reach out on Upwork and I'll take a look.
Top comments (0)