Last month, I published a benchmark showing a 1.125× speedup from block-residual caching on 4-bit FLUX.

The main lesson was not the multiplier. It was that my original quality metrics had been measuring the wrong thing, and that acceleration claims often combine speed, trajectory preservation, and perceptual quality into one number.

For the follow-up, I chose a stricter target: real-time autoregressive diffusion video on an Apple M5 Max, with the definition of "real time" frozen before results were visible.

The tested configuration did not meet that target. The fastest claim-eligible result was 1.418 native generated frames per second, compared with a 16 FPS target. That is an 11.28× gap.

I am publishing the result because the measured bottleneck, one systems improvement, and two rejected hypotheses are useful even without a real-time result.