Early Memory Selection for Balanced Adam
Alberto Fernández-Hernández, Cristian Pérez-Corral, Jose I. Mestre, Manuel F. Dolz, Enrique S. Quintana-Ortí
- Published
- Oct 6, 2026 — 16:22 UTC
Problem
This work addresses a gap in the selection of the shared memory parameter B2_1 = B2_2 = B2 in the Adam optimizer, which is critical for balancing convergence speed and stability. The authors propose a method to dynamically select this parameter based on early training data, which is particularly relevant given the lack of prior literature on adaptive selection of B2 in Adam.
Method
The proposed method, Early Memory Selection for Adam, involves selecting the shared memory parameter B2 from a short pilot training phase consisting of 200 updates. The approach utilizes a local model of Adam's normalized direction to balance sampling variability against the delay introduced by averaging past gradients. A cubic memory rule is employed, with two coefficients estimated from gradient probes. Specifically, the method uses sixteen probe gradients collected at four checkpoints during the pilot training phase. The estimator jointly considers both the numerator and denominator to maintain covariance, enhancing the robustness of the parameter selection process.
Results
The results demonstrate significant improvements in validation gap reduction compared to baseline configurations:
- Mean Relative Validation Gap Reduction: 40.7% compared to a shared B2 = 0.95
- Worst-Quarter Mean Gap Reduction: 44.3% compared to a shared B2 = 0.95
- Mean Gap Reduction: 32.3% compared to the best constant B2 across eleven different workloads.
These results indicate that the proposed method effectively enhances the performance of the Adam optimizer across diverse scenarios.
Limitations
The authors do not report any limitations in their work, and no obvious limitations are identified in the provided text.
Why it matters
The implications of this research are significant for downstream work in optimization algorithms, particularly in scenarios where adaptive parameter selection can lead to improved convergence rates and model performance. By addressing the selection of the shared memory parameter in Adam, this work opens avenues for further exploration of adaptive methods in other optimization contexts, potentially leading to more efficient training processes in machine learning.
By Turing Wire Research Desk · Oct 6, 2026 · How we work →
Summarised from the paper by the Turing Wire Research Desk. The full paper has the complete methods and results.
Source: arXiv cs.AI
