Adam drifts far from natural gradient descent under ill-conditioning, paper finds
Adam is the standard optimizer in deep learning, but its geometric relationship to natural gradient descent (NGD) has unresolved questions. A new arXiv paper tackles them by studying Adam's full update rule, momentum included, as a diagonal empirical Fisher approximation. In that framing Adam is subject to three sources of error: diagonal truncation, empirical label substitution, and temporal lag.
To quantify how far Adam wanders from the real thing, the authors use a scale-invariant metric called gamma(Delta theta) and measure Adam's geometric deviation from true NGD across four loss landscapes: well-conditioned linear regression, ill-conditioned linear regression, logistic regression, and a non-convex small neural network.
The central finding is that Adam's geometric trajectory is context-dependent. Deviation from NGD stays low in well-conditioned settings but rises significantly under ill-conditioning, reaching misalignments of about 10^3 in the neural network. Higher geometric drift correlates with slower initial optimization, but it does not degrade final objective minimization: Adam consistently reaches low loss.
The paper also compares two Fisher variants. The improved empirical Fisher (iEF) tracks more stable paths than the standard empirical Fisher (EF), which frequently oscillates or diverges.
From this the authors draw a cautious conclusion. Their results suggest that Adam's practical optimization power may stem from a balance of structural approximation errors and momentum smoothing, rather than from close tracking of the natural gradient path.
Key facts
- The paper treats Adam's full update rule, including momentum, as a diagonal empirical Fisher approximation with three error sources: diagonal truncation, empirical label substitution and temporal lag.
- Deviation from true natural gradient descent is measured with the scale-invariant gamma(Delta theta) metric across four loss landscapes: well-conditioned and ill-conditioned linear regression, logistic regression, and a non-convex small neural network.
- Deviation stays low when the problem is well-conditioned but rises significantly under ill-conditioning, reaching misalignments of about 10^3 in the neural network.
- Greater geometric drift goes with slower initial optimization but does not hurt final objective minimization; Adam consistently reaches low loss.
- The improved empirical Fisher (iEF) follows more stable paths than the standard empirical Fisher (EF), which frequently oscillates or diverges.
Why it matters
Adam is the standard optimizer in deep learning, and a common way to explain it is as an approximation to natural gradient descent. The paper asks how good that approximation really is, and its answer is: it depends on the problem. When conditioning is good, Adam stays close to NGD. When conditioning is poor, it departs a long way, with misalignments of about 10^3 in the neural network case. Yet it still ends up at low loss. That pulls apart two ideas that are often run together: following the natural gradient path closely, and optimizing well in practice.
Who it affects
The work is aimed at researchers who study optimizer theory and the geometry of training, especially those who use the Fisher information or natural gradient picture to reason about Adam. It also bears on anyone comparing empirical Fisher variants, since the iEF is reported to be steadier than the standard EF, which frequently oscillates or diverges.
How to use it
The abstract does not claim a practical recommendation or a new optimizer, so there is nothing to switch on. Its value is as a lens: when analysing Adam, include momentum and the three approximation errors (diagonal truncation, empirical label substitution, temporal lag) instead of treating it as a plain diagonal Fisher method, and expect larger departures from NGD on ill-conditioned problems. The gamma(Delta theta) metric is the tool the authors use to measure that departure.
How solid is it
The material here is the paper's abstract, so the detail behind the experiments is not visible. The claims are stated by the authors, and the final explanation is hedged: they say the results suggest Adam's power may stem from a balance of approximation errors and momentum smoothing. The abstract names no authors or institutions. The size and architecture of the small neural network and the datasets are not specified, and no numeric values are given for the deviation in the well-conditioned settings, for the linear or logistic regression cases, or for the slowdown in initial optimization.
Risks and caveats
Four loss landscapes, one of them a small neural network, are a narrow base from which to generalise to large-scale deep learning. The roughly 10^3 misalignment is a value of the gamma metric; the abstract does not say whether it is a mean, a maximum or a final value, and does not describe it as a percentage or a ratio to a baseline. The conclusion about momentum smoothing is a suggestion, not a demonstrated mechanism. The abstract mentions no comparison with optimizers other than Adam, EF and iEF.
“Our results suggest Adam's practical optimization power may stem from a balance of structural approximation errors and momentum smoothing rather than close tracking of the natural gradient path.”
— Abstract of arXiv paper 2610.00004