A small neural network learns to add numbers on a clock with 59 hours. It trains on 40% of the 3481 possible sums and never sees the rest. It memorises its training sums within a few hundred steps, but stays at chance on the unseen ones. Thousands of steps later, it suddenly gets them right. That delayed jump is grokking.
Pick two numbers, or click a cell in the grid. The bars use a log scale, so each gridline is 100,000 times less likely than the one above. Before grokking the unlikely answers are noise. After, they rise and fall in waves that peak at the right answer.
The network gives each number a vector. Early on, those vectors mix every frequency. After grokking, a few frequencies (bright bars) carry most of the energy: the network has learned to represent each number as waves.
Each number, projected onto the strongest frequency. Once grokked they form a circle, like a clock face. Adding b becomes a rotation by b steps, which works for every pair, seen or not.
Training on 1,392 of 3,481 sums, full batch, AdamW. Try weight decay 0, then restart: it memorises and stays there. Weight decay keeps squeezing the weights, and the general rule needs far less weight than a lookup table.