We had a large table - hundreds of gigabytes - that needed to be fully exported and processed in batches. The pipeline fetched rows from Postgres in batches of 2000, processed them, and pushed results downstream.

It was slow. Much slower than expected. The bottleneck wasn't the processing. It was Postgres. Specifically, how we were reading from it.

How Postgres Stores Data

Before getting to the fix, you need to understand three things: pages, slots, and sequence columns.

Pages are fixed 8KB chunks on disk. Everything Postgres stores lives inside pages.