AdamW optimizer for bfloat16 models in pytorch.

Alex Rogozhnikov

Last update: Nov 20, 2022

Related tags

Deep Learning adamw_bfloat16

Overview

_{Image source}

AdamW optimizer for bfloat16 models in pytorch.

Bfloat16 is currently an optimal tradeoff between range and relative error for deep networks.
Bfloat16 can be used quite efficiently on Nvidia GPUs with Ampere architecture (A100, A10, A30, RTX3090...)

However, neither AMP in pytorch is ready for bfloat16, nor optimizers.

If you just convert all weights and inputs to bfloat16, you're likely to run into an issue of stale weights: updates are too small to modify bfloat16 weight (see gopher paper, section C2 for a large-scale example).

There are two possible remedies:

keep weights in float32 (precise) and bfloat16 (approximate)
keep weights in bfloat16, and keep correction term in bfloat16

As recent study has shown, both options are completely competitive in quality to float32 training.

Usage

Install:

pip install git+https://github.com/arogozhnikov/adamw_bfloat16.git

Use as a drop-in replacement for pytorch's AdamW:

import torch
from adamw_bfloat16 import LR, AdamW_BF16
model = model.to(torch.bfloat16)

# default preheat and decay
optimizer = AdamW_BF16(model.parameters())

# configure LR schedule. Use built-in scheduling opportunity
optimizer = AdamW_BF16(model.parameters(), lr_function=LR(lr=1e-4, preheat_steps=5000, decay_power=-0.25))

Releases(v0.1.0)

v0.1.0(Dec 14, 2021)

Initial implementation of AdamW for pytorch supports cuda graphs and has a built-in mechanism for control of learning rate, because external are unlikely to make a friendship with cuda graphs
Source code(tar.gz)
Source code(zip)

AdamW optimizer for bfloat16 models in pytorch.

Related tags

Overview

AdamW optimizer for bfloat16 models in pytorch.

Usage

You might also like...

Storage-optimizer - Identify potintial optimizations on the cloud storage accounts

PyTorch implementation and pretrained models for XCiT models. See XCiT: Cross-Covariance Image Transformer

Objective of the repository is to learn and build machine learning models using Pytorch. 30DaysofML Using Pytorch

Pretrained SOTA Deep Learning models, callbacks and more for research and production with PyTorch Lightning and PyTorch

A bunch of random PyTorch models using PyTorch's C++ frontend

PyTorch-LIT is the Lite Inference Toolkit (LIT) for PyTorch which focuses on easy and fast inference of large models on end-devices.

Pytorch-diffusion - A basic PyTorch implementation of 'Denoising Diffusion Probabilistic Models'

pyhsmm - library for approximate unsupervised inference in Bayesian Hidden Markov Models (HMMs) and explicit-duration Hidden semi-Markov Models (HSMMs), focusing on the Bayesian Nonparametric extensions, the HDP-HMM and HDP-HSMM, mostly with weak-limit approximations.

Releases(v0.1.0)

v0.1.0(Dec 14, 2021)

Owner

Alex Rogozhnikov

ESGD-M - A stochastic non-convex second order optimizer, suitable for training deep learning models, for PyTorch

PyTorch implementation DRO: Deep Recurrent Optimizer for Structure-from-Motion

A mini library for Policy Gradients with Parameter-based Exploration, with reference implementation of the ClipUp optimizer from NNAISENSE.

Ranger deep learning optimizer rewrite to use newest components

auto-tuning momentum SGD optimizer

Ranger - a synergistic optimizer using RAdam (Rectified Adam), Gradient Centralization and LookAhead in one codebase

Apollo optimizer in tensorflow

This is an implementation of Googles Yogi-Optimizer in Keras (tf.keras)

DeepOBS: A Deep Learning Optimizer Benchmark Suite

An Implicit Function Theorem (IFT) optimizer for bi-level optimizations