Back to Articles

Hierarchical pooling already lets you cut ColBERT storage in half without hurting retrieval. We now show that STE-based regularization, the same trick we used to fix MUVERA/SMVE, pushes this much further: 99.4% retention at 5× compression, without hurting full-token performance. The key insight is that the same as for MUVERA/SMVE: the best way to achieve compressability is to directly train for it.

Table of Contents

Hierarchical Pooling: The Quiet Workhorse

A Regularization That Keeps on Giving