Batch Normalization: The Tiny Trick That Changed Deep Learning

If you’ve ever dipped your toes into deep learning—especially convolutional neural networks (CNNs) for computer vision—you’ve probably encountered Batch Normalization (BatchNorm).

At first glance, BatchNorm seems almost too simple: take the activations, subtract the mean, divide by the standard deviation, then scale and shift them. Yet this small mathematical operation has had a profound impact on how deep neural networks are trained.

Let’s unpack the story of BatchNorm—from its original motivation as a way to stabilize training to modern insights that reveal why it works so well.

Why Normalize Inputs in the First Place?