A short field note. I had to crawl every item in a public registry (~4 million of them) in about a day. One idea did most of the work.

The problem

Split the sorted keyspace into N ranges, run one worker per range, done. That's the plan. It falls apart because keys aren't uniformly distributed. Cut the alphabet into even chunks and one chunk swallows all the dense prefixes while the rest finish in minutes and sit idle. My first cut had a single shard holding 71% of the work.

And your crawl is only as fast as the slowest shard. So that skew throws away almost all the parallelism: a 56-way fan-out that finishes no faster than a 3-way one. I didn't need equal alphabet width. I needed equal item count per range.

The trick: your source already ranks keys