Despite the tremendous success of Stochastic Gradient Descent (SGD) algorithm\nin deep learning, little is known about how SGD finds generalizable solutions\nin the high-dimensional weight space. By analyzing the learning dynamics and\nloss function landscape, we discover a robust inverse relation between the\nweight variance and the landscape flatness (inverse of curvature) for all\nSGD-based learning algorithms. To explain the inverse variance-flatness\nrelation, we develop a random landscape theory, which shows that the SGD noise\nstrength (effective temperature) depends inversely on the landscape flatness.\nOur study indicates that SGD attains a self-tuned landscape-dependent annealing\nstrategy to find generalizable solutions at the flat minima of the landscape.\nFinally, we demonstrate how these new theoretical insights lead to more\nefficient algorithms, e.g., for avoiding catastrophic forgetting.\n