Part of our Inference Engineering Masterclass pod yesterday involved a spicy discussion about Megakernels:megakernels are deadwhy are megakernels useful? you spend two months writing a kernel to save time on launch overhead and poor inter-kernel overlap. you had PDL but then people said it wasn't perfect, that you could still get some marginal gains due to straggler CTAs and therefore- wait, sorry, I forgot, Rubin fixes that (kernel two needs 10 CTAs and kernel one has seven finished and three straggling, kernel two launches seven of its CTAs). given a long enough timeline, it all evens out. no serious inference provider is using a 67k loc hand-fused forward pass kernel in production, and the teams doing that are doing so out of pure research. dead.ali@waterloo_interntwo weeks ago i went on @swyx's pod and said some things that i... should not have said.

a lot has happened since then, i owe you all an apology.

i'm sorry that i was right about every single thing.

a) re megakernels are dead

why are megakernels useful? you spend two monthsLatent.Space @latentspacepodThe Inference Engineering Masterclass: 10x faster models, quantization, speculative decoding, Rubin, & self-optimizing AI https://t.co/uRYIWWDebj