Modular arithmetic·(a + b) mod 59

Grokking

Step
0
Phase
Memorising

A small neural network learns to add numbers on a clock with 59 hours. It trains on 40% of the 3481 possible sums and never sees the rest. It memorises its training sums within a few hundred steps, but stays at chance on the unseen ones. Thousands of steps later, it suddenly gets them right. That delayed jump is grokking.

01Accuracy
TrainUnseen
0%25%50%75%100%1101001,000
02Loss
TrainUnseen
1011e-11e-21e-31101001,000
03Weight norm
05101101001,000
04Telemetry
Train acc
-%
Unseen acc
-%
Weight norm
-
Top 5 freqs
-%
Fitting the training pairs.
05Every sum: row a, column b
Train, rightTrain, wrongUnseen, rightUnseen, wrong
06Ask the network
+mod 59 =?
Confidence
-
Right answer
14 ✗
Trained on it
Never seen

Pick two numbers, or click a cell in the grid. The bars use a log scale, so each gridline is 100,000 times less likely than the one above. Before grokking the unlikely answers are noise. After, they rise and fall in waves that peak at the right answer.

07Fourier spectrum of the number embeddings

The network gives each number a vector. Early on, those vectors mix every frequency. After grokking, a few frequencies (bright bars) carry most of the energy: the network has learned to represent each number as waves.

08Embeddings

Each number, projected onto the strongest frequency. Once grokked they form a circle, like a clock face. Adding b becomes a rotation by b steps, which works for every pair, seen or not.

09Control deck

Training on 1,392 of 3,481 sums, full batch, AdamW. Try weight decay 0, then restart: it memorises and stays there. Weight decay keeps squeezing the weights, and the general rule needs far less weight than a lookup table.