ovr.news

Solutions that work, including long-horizon plans with outcomes

AdaGC Stabilizes Large Language Model Training

arxiv.org · 21 July 2026
AdaGC Stabilizes Large Language Model Training
Photo: arxiv.org
Read on arxiv.org

Researchers propose AdaGC, a new adaptive gradient clipping scheme, to improve the stability of large language model pretraining. They observed that loss spikes, disruptions in training, arise from multiple factors including data issues and computational errors.

AdaGC addresses this by bounding gradient norms based on a tensor’s historical clipped values. The team designed AdaGC to work with any optimization algorithm and minimize additional memory use and communication demands, especially during distributed training.

Testing on Llama-2 7B, Mixtral 8x1B, and ERNIE 10B-A1.4B models showed AdaGC eliminated training instabilities and improved downstream accuracy compared to GlobalGC by up to 2.48%. The code is publicly available on GitHub.

Surfaced by the Solutions lens — one of the vital signs ovr.news reads.

How we evaluated this
AI summary

read the original for the full story — Read on arxiv.org . How we work →

Why are you reporting this article?

Why are you reporting this article?