Storia: LayerNorm vs BatchNorm: why Transformers normalize per token, not per batch — Warptech Lab News