I ran a scraping platform that processed millions of pages a day at roughly 95% extraction success, around three seconds per page. The fetch-and-parse code, the part everyone thinks of as "the scraper", was a tiny slice of the whole thing. The years went into everything around it.
Here's what actually broke, more or less in the order it broke.
The queue with no manners
Crawling is bursty in a way that surprises you the first time. One category page fans out into a few hundred product URLs. A sitemap refresh dumps half a million at once. Meanwhile your parsers chew through pages at a steady rate that doesn't care about your ambitions.
Our first queue was effectively unbounded. It absorbed every burst happily, Redis memory climbed for two days, and then the whole thing fell over at once instead of slowing down gracefully. Lesson: if your queue can't say no, it's not a queue, it's a landfill. Bound it, and make producers block or shed when it's full. A crawler that pauses discovery for an hour is a non-event. A crawler that OOMs the broker is a weekend.






