A small convolutional network learns to read handwritten digits from just 1,000 examples. It starts with deliberately oversized weights, so it memorises its training images within about 2,000 steps while reading only about 78% of unseen digits correctly. Then it stalls. Weight decay keeps shrinking the weights until the memorised lookup table no longer fits, and the network switches to features that work on digits it has never seen.
Draw a digit, or click one in the strip below. The bars update live as it trains. It learned from only 1,000 examples, so unusual styles trip it up: draw 4s open at the top, as most MNIST 4s are.
Loading digits...
Test digits the network never trains on. Red frames are ones it reads wrong. Watch the red thin out after the plateau. Click one to see its probabilities.
Mini-batches of 100, AdamW. Training pauses at step 8,000. Try 1x starting weights and restart: it learns unseen digits straight away, with no plateau. The oversized start is what makes it memorise first.
Digits from the MNIST database by Yann LeCun, Corinna Cortes and Christopher J.C. Burges, used under CC BY-SA 3.0. Shrunk to 14x14 pixels. The oversized starting weights follow Liu and others, 'Omnigrok: grokking beyond algorithmic data', 2022.