TTFT, TPOT, end-to-end latency—learn what inference latency actually measures, where the time goes in production pipelines, and when skipping the call beats speeding it up.