Storia: Inference Optimization for the Rest of Us — KV Cache, Quantization, and Latency Tradeoffs — Warptech Lab News