Grokking and Phase Transitions: A Pedagogical Approach to the Statistical Physics of Machine Learning
DOI:
https://doi.org/10.1590/SciELOPreprints.16570Keywords:
Grokking, Statistical Physics, Neural Networks, Phase Transition, Deep Learning, Machine LearningAbstract
The phenomenon of grokking in artificial neural networks describes the counter-intuitive situation in which a model, after completely overfitting the training data for a long period, suddenly achieves perfect generalization on unseen test data. In this work, we present a pedagogical review of this phenomenon from the perspective of Statistical Physics, aimed at undergraduate education in physics and related areas. We demonstrate how the learning process can be mapped as a first-order dynamic phase transition, where the validation accuracy plays the role of an order parameter, weight decay acts as a restrictive harmonic potential, and the random weight initialization combined with the chaotic/stochastic dynamic of the optimizer under finite learning rates emulate an effective temperature whose inversion favors low-energy ordered states. To make the concepts concrete, we provide reproducible Python codes for training and analyzing Multilayer Perceptron (MLP) and Transformer networks applied to three distinct experiments: modular arithmetic, the sparse parity problem, and modular energy conservation. We show that generalization emerges from discovering hidden symmetries in the dataset, manifested by the periodic alignment of weights in the form of circular trigonometric representations (analyzed via Fourier Transform) and by the concentric spatial organization of degenerate energy states in the embedding space.
Downloads
Submitted
Posted
How to Cite
Section
Copyright (c) 2026 João Pedro Sansão

This work is licensed under a Creative Commons Attribution 4.0 International License.
Plaudit
Data statement
-
The research data is available in one or more data repository(ies)


