Storia: Hardware-aware framework accelerates large language models without additional training — Warptech Lab News