Back to Articles
Hierarchical pooling already lets you cut ColBERT storage in half without hurting retrieval. We now show that STE-based regularization, the same trick we used to fix MUVERA/SMVE, pushes this much further: 99.4% retention at 5× compression, without hurting full-token performance. The key insight is that the same as for MUVERA/SMVE: the best way to achieve compressability is to directly train for it.
Table of Contents
Hierarchical Pooling: The Quiet Workhorse
A Regularization That Keeps on Giving








