imbalanced-DL: Deep Imbalanced Learning in Python

NTUCSIE CLLab

Last update: Dec 28, 2022

Related tags

Deep Learning imbalanced-DL

Overview

imbalanced-DL: Deep Imbalanced Learning in Python

Overview

imbalanced-DL (imported as imbalanceddl) is a Python package designed to make deep imbalanced learning easier for researchers and real-world users. From our experiences, we observe that to tackcle deep imbalanced learning, there is a need for a strategy. That is, we may not just address this problem with one single model or approach. Thus in this package, we seek to provide several strategies for deep imbalanced learning. The package not only implements several popular deep imbalanced learning strategies, but also provides benchmark results on several image classification tasks. Futhermore, this package provides an interface for implementing more datasets and strategies.

Strategy

We provide some baseline strategies as well as some state-of-the-are strategies in this package as the following:

Empirical Risk Minimization (baseline strategy)
Reweighting with Class Balance (CB) Loss
Deferred Re-Weighting (DRW)
Label Distribution Aware Margin (LDAM) Loss with DRW
Mixup with DRW
Remix with DRW

Environments

This package is tested on Linux OS.
You are suggested to use a different virtual environment so as to avoid package dependency issue.
For Pyenv & Virtualenv users, you can follow the below steps to create a new virtual environment or you can also skip this step.

Pyenv & Virtualenv (Optinal)

For dependency isolation, it's better to create another virtual environment for usage.
The following will be the demo for creating and managing virtual environment.
Install pyenv & virtualenv first.
pyenv virtualenv [version] [virtualenv_name]
- For example, if you'd like to use python 3.6.8, you can do: pyenv virtualenv 3.6.8 TestEnv
mkdir [dir_name]
cd [dir_name]
pyenv local [virtualenv_name]
Then, you will have a new (clean) python virtual environment for the package installation.

Installation

Basic Requirement

Python >= 3.6

git clone https://github.com/ntucllab/imbalanced-DL.git
cd imbalanceddl
python -m pip install -r requirements.txt
python setup.py install

Usage

We highlight three key features of imbalanced-DL as the following:

(0) Imbalanced Dataset:

We support 5 benchmark image datasets for deep imbalanced learing.
To create and ImbalancedDataset object, you will need to provide a config_file as well as the dataset name you would like to use.
Specifically, inside the config_file, you will need to specify three key parameters for creating imbalanced dataset.
- imb_type: you can choose from exp (long-tailed imbalance) or step imbalanced type.
- imb_ratio: you can specify the imbalanceness of your data, typically researchers choose 0.1 or 0.01.
- dataset_name: you can specify 5 benchmark image datasets we provide, or you can implement your own dataset.
- For an example of the config_file, you can see example/config.
To contruct your own dataset, you should inherit from BaseDataset, and you can follow torchvision.datasets.ImageFolder to construct your dataset in PyTorch format.

from imbalanceddl.dataset.imbalance_dataset import ImbalancedDataset

# specify the dataset name
imbalance_dataset = ImbalancedDataset(config, dataset_name=config.dataset)

(1) Strategy Trainer:

We support 6 different strategies for deep imbalance learning, and you can either choose to train from scratch, or evaluate with the best model after training. To evaluate with the best model, you can get more in-depth metrics such as per class accuracy for further evaluation on the performance of the selected strategy. We provide one trained model in example/checkpoint_cifar10.
For each strategy trainer, it is associated with a config_file, ImbalancedDataset object, model, and strategy_name.
Specifically, the config_file will provide some training parameters, where the default settings for reproducing benchmark result can be found in example/config. You can also set these training parameters based on your own need.
For model, we currently provide resnet32 and resnet18 for reproducing the benchmark results.
We provide a build_trainer() function to return the specified trainer as the following.

from imbalanceddl.strategy.build_trainer import build_trainer

# specify the strategy
trainer = build_trainer(config,
                        imbalance_dataset,
                        model=model,
                        strategy=config.strategy)
# train from scratch
trainer.do_train_val()

# Evaluate with best model
trainer.eval_best_model()

Or you can also just select the specific strategy you would like to use as:

from imbalanceddl.strategy import LDAMDRWTrainer

# pick the trainer
trainer = LDAMDRWTrainer(config,
                         imbalance_dataset,
                         model=model,
                         strategy=config.strategy)

# train from scratch
trainer.do_train_val()

# Evaluate with best model
trainer.eval_best_model()

To construct your own strategy trainer, you need to inherit from Trainer class, where in your own strategy you will have to implement get_criterion() and train_one_epoch() method. After this you can choose whether to add your strategy to build_trainer() function or you can just use it as the above demonstration.

(2) Benchmark research environment:

To conduct deep imbalanced learning research, we provide example codes for training with different strategies, and provide benchmark results on five image datasets. To quickly start training CIFAR-10 with ERM strategy, you can do:

cd example
python main.py --gpu 0 --seed 1126 --c config/config_cifar10.yaml --strategy ERM

Following the example code, you can not only get results from baseline training as well as state-of-the-art performance such as LDAM or Remix, but also use this environment to develop your own algorithm / strategy. Feel free to add your own strategy into this package.
For more information about example and usage, please see the Example README

Benchmark Results

We provide benchmark results on 5 image datasets, including CIFAR-10, CIFAR-100, CINIC-10, SVHN, and Tiny-ImageNet. We follow standard procedure to generate imbalanced training dataset for these 5 datasets, and provide their top 1 validation accuracy results for research benchmark. For example, below you can see the result table of Long-tailed Imbalanced CIFAR-10 trained on different strategies. For more detailed benchmark results, please see example/README.md.

Long-tailed Imbalanced CIFAR-10

`imb_type`	`imb_factor`	Model	Strategy	Validation Top 1
long-tailed	100	ResNet32	ERM	71.23
long-tailed	100	ResNet32	DRW	75.08
long-tailed	100	ResNet32	LDAM-DRW	77.75
long-tailed	100	ResNet32	Mixup-DRW	82.11
long-tailed	100	ResNet32	Remix-DRW	81.82

Test

python -m unittest -v

Contact

If you have any question, please don't hesitate to email [email protected]. Thanks !

Acknowledgement

The authors thank members of the Computational Learning Lab at National Taiwan University for valuable discussions and various contributions to making this package better.

Comments

Integrated m2mmethod
I added M2m method in our package. Here is the list of modification:

Add m2m data loader for 5 datasets: CIFAR10, CIFAR100, SVHN10, CINIC10, TINY200

Add _m2m strategy

Update READ.me files

Separate a new config file instead of putting it in main file

Add more functions support for m2m method
opened by maitanha 0

Ivy is a templated deep learning framework which maximizes the portability of deep learning codebases.

Ivy is a templated deep learning framework which maximizes the portability of deep learning codebases. Ivy wraps the functional APIs of existing frameworks. Framework-agnostic functions, libraries and layers can then be written using Ivy, with simultaneous support for all frameworks. Ivy currently supports Jax, TensorFlow, PyTorch, MXNet and Numpy. Check out the docs for more info!

8.2k Jan 2, 2023

Deep learning (neural network) based remote photoplethysmography: how to extract pulse signal from video using deep learning tools

Deep-rPPG: Camera-based pulse estimation using deep learning tools Deep learning (neural network) based remote photoplethysmography: how to extract pu

138 Dec 17, 2022

deep-table implements various state-of-the-art deep learning and self-supervised learning algorithms for tabular data using PyTorch.

63 Oct 17, 2022

Time-series-deep-learning - Developing Deep learning LSTM, BiLSTM models, and NeuralProphet for multi-step time-series forecasting of stock price.

Stock Price Prediction Using Deep Learning Univariate Time Series Predicting stock price using historical data of a company using Neural networks for

7 Nov 27, 2022

Deep Learning: Architectures & Methods Project: Deep Learning for Audio Super-Resolution

Deep Learning: Architectures & Methods Project: Deep Learning for Audio Super-Resolution Figure: Example visualization of the method and baseline as a

16 Dec 23, 2022

PyTorch implementation of the Deep SLDA method from our CVPRW-2020 paper "Lifelong Machine Learning with Deep Streaming Linear Discriminant Analysis"

Lifelong Machine Learning with Deep Streaming Linear Discriminant Analysis This is a PyTorch implementation of the Deep Streaming Linear Discriminant

41 Dec 25, 2022

Deep Image Search is an AI-based image search engine that includes deep transfor learning features Extraction and tree-based vectorized search.

Deep Image Search - AI-Based Image Search Engine Deep Image Search is an AI-based image search engine that includes deep transfer learning features Ex

139 Jan 1, 2023

H2O is an Open Source, Distributed, Fast & Scalable Machine Learning Platform: Deep Learning, Gradient Boosting (GBM) & XGBoost, Random Forest, Generalized Linear Modeling (GLM with Elastic Net), K-Means, PCA, Generalized Additive Models (GAM), RuleFit, Support Vector Machine (SVM), Stacked Ensembles, Automatic Machine Learning (AutoML), etc.

H2O H2O is an in-memory platform for distributed, scalable machine learning. H2O uses familiar interfaces like R, Python, Scala, Java, JSON and the Fl

6.1k Jan 5, 2023

Machine Learning From Scratch. Bare bones NumPy implementations of machine learning models and algorithms with a focus on accessibility. Aims to cover everything from linear regression to deep learning.

Machine Learning From Scratch About Python implementations of some of the fundamental Machine Learning models and algorithms from scratch. The purpose

21.8k Jan 9, 2023

imbalanced-DL: Deep Imbalanced Learning in Python

Related tags

Overview

imbalanced-DL: Deep Imbalanced Learning in Python

Overview

Strategy

Environments

Pyenv & Virtualenv (Optinal)

Installation

Basic Requirement

Usage

Benchmark Results

Test

Contact

Acknowledgement

You might also like...

Ivy is a templated deep learning framework which maximizes the portability of deep learning codebases.

Deep learning (neural network) based remote photoplethysmography: how to extract pulse signal from video using deep learning tools

deep-table implements various state-of-the-art deep learning and self-supervised learning algorithms for tabular data using PyTorch.

Time-series-deep-learning - Developing Deep learning LSTM, BiLSTM models, and NeuralProphet for multi-step time-series forecasting of stock price.

Deep Learning: Architectures & Methods Project: Deep Learning for Audio Super-Resolution

PyTorch implementation of the Deep SLDA method from our CVPRW-2020 paper "Lifelong Machine Learning with Deep Streaming Linear Discriminant Analysis"

Deep Image Search is an AI-based image search engine that includes deep transfor learning features Extraction and tree-based vectorized search.

Machine Learning From Scratch. Bare bones NumPy implementations of machine learning models and algorithms with a focus on accessibility. Aims to cover everything from linear regression to deep learning.

Comments

Integrated m2mmethod

Owner

NTUCSIE CLLab

[ICML 2021, Long Talk] Delving into Deep Imbalanced Regression

A Pytorch implementation of CVPR 2021 paper "RSG: A Simple but Effective Module for Learning Imbalanced Datasets"

[NeurIPS 2021] “Improving Contrastive Learning on Imbalanced Data via Open-World Sampling”,

The repo contains the code of the ACL2020 paper `Dice Loss for Data-imbalanced NLP Tasks`

MetaBalance: High-Performance Neural Networks for Class-Imbalanced Data

A (PyTorch) imbalanced dataset sampler for oversampling low frequent classes and undersampling high frequent ones.

Official implementation of Influence-balanced Loss for Imbalanced Visual Classification in PyTorch.

Imbalanced Gradients: A Subtle Cause of Overestimated Adversarial Robustness

BESS: Balanced Evolutionary Semi-Stacking for Disease Detection via Partially Labeled Imbalanced Tongue Data

FTIR-Deep Learning - FTIR Deep Learning With Python