Local LLM inference has an expensive habit:

It recomputes prefixes it has already seen.

A system prompt.

A reused RAG document.

A few-shot block.