Revitalizing CNN Attention via Transformers in Self-Supervised Visual Representation Learning

ChongjianGE

Last update: Dec 2, 2022

Related tags

Deep Learning CARE

Overview

Revitalizing CNN Attention via Transformers in Self-Supervised Visual Representation Learning

This repository is the official implementation of CARE.

Updates

(09/10/2021) Our paper is accepted by NeurIPS 2021.

Requirements

To install requirements:

conda create -n care python=3.6
conda install pytorch==1.7.1 torchvision==0.8.2 torchaudio==0.7.2 cudatoolkit=10.1 -c pytorch
pip install tensorboard
pip install ipdb
pip install einops
pip install loguru
pip install pyarrow==3.0.0
pip install tqdm

📋 Pytorch>=1.6 is needed for runing the code.

Data Preparation

Prepare the ImageNet data in {data_path}/train.lmdb and {data_path}/val.lmdb

Relpace the original data path in care/data/dataset_lmdb (Line7 and Line40) with your new {data_path}.

📋 Note that we use the lmdb file to speed-up the data-processing procedure.

Training

Before training the ResNet-50 (100 epoch) in the paper, run this command first to add your PYTHONPATH:

export PYTHONPATH=$PYTHONPATH:{your_code_path}/care/
export PYTHONPATH=$PYTHONPATH:{your_code_path}/care/care/

Then run the training code via:

bash run_train.sh      #(The training script is used for trianing CARE with 8 gpus)
bash single_gpu_train.sh    #(We also provide the script for trainig CARE with only one gpu)

📋 The training script is used to do unsupervised pre-training of a ResNet-50 model on ImageNet in an 8-gpu machine

using -b to specify batch_size, e.g., -b 128

using -d to specify gpu_id for training, e.g., -d 0-7

using --log_path to specify the main folder for saving experimental results.

using --experiment-name to specify the folder for saving training outputs.

The code base also supports for training other backbones (e.g., ResNet101 and ResNet152) with different training schedules (e.g., 200, 400 and 800 epochs).

Evaluation

Before start the evaluation, run this command first to add your PYTHONPATH:

export PYTHONPATH=$PYTHONPATH:{your_code_path}/care/
export PYTHONPATH=$PYTHONPATH:{your_code_path}/care/care/

Then, to evaluate the pre-trained model (e.g., ResNet50-100epoch) on ImageNet, run:

bash run_val.sh      #(The training script is used for evaluating CARE with 8 gpus)
bash debug_val.sh    #(We also provide the script for evaluating CARE with only one gpu)

📋 The training script is used to do the supervised linear evaluation of a ResNet-50 model on ImageNet in an 8-gpu machine

using -b to specify batch_size, e.g., -b 128

using -d to specify gpu_id for training, e.g., -d 0-7

Modifying --log_path according to your own config.

Modifying --experiment-name according to your own config.

Pre-trained Models

We here provide some pre-trained models in the [shared folder]:

Here are some examples.

[ResNet-50 100epoch] trained on ImageNet using ResNet-50 with 100 epochs.
[ResNet-50 200epoch] trained on ImageNet using ResNet-50 with 200 epochs.
[ResNet-50 400epoch] trained on ImageNet using ResNet-50 with 400 epochs.

More models are provided in the following model zoo part.

📋 We will provide more pretrained models in the future.

Model Zoo

Our model achieves the following performance on :

Self-supervised learning on image classifications.

Method	Backbone	epoch	Top-1	Top-5	pretrained model	linear evaluation model
CARE	ResNet50	100	72.02%	90.02%	[pretrained] (wip)	[linear_model] (wip)
CARE	ResNet50	200	73.78%	91.50%	[pretrained] (wip)	[linear_model] (wip)
CARE	ResNet50	400	74.68%	91.97%	[pretrained] (wip)	[linear_model] (wip)
CARE	ResNet50	800	75.56%	92.32%	[pretrained] (wip)	[linear_model] (wip)
CARE	ResNet50(2x)	100	73.51%	91.66%	[pretrained] (wip)	[linear_model] (wip)
CARE	ResNet50(2x)	200	75.00%	92.22%	[pretrained] (wip)	[linear_model] (wip)
CARE	ResNet50(2x)	400	76.48%	92.99%	[pretrained] (wip)	[linear_model] (wip)
CARE	ResNet50(2x)	800	77.04%	93.22%	[pretrained] (wip)	[linear_model] (wip)
CARE	ResNet101	100	73.54%	91.63%	[pretrained] (wip)	[linear_model] (wip)
CARE	ResNet101	200	75.89%	92.70%	[pretrained] (wip)	[linear_model] (wip)
CARE	ResNet101	400	76.85%	93.31%	[pretrained] (wip)	[linear_model] (wip)
CARE	ResNet101	800	77.23%	93.52%	[pretrained] (wip)	[linear_model] (wip)
CARE	ResNet152	100	74.59%	92.09%	[pretrained] (wip)	[linear_model] (wip)
CARE	ResNet152	200	76.58%	93.63%	[pretrained] (wip)	[linear_model] (wip)
CARE	ResNet152	400	77.40%	93.63%	[pretrained] (wip)	[linear_model] (wip)
CARE	ResNet152	800	78.11%	93.81%	[pretrained] (wip)	[linear_model] (wip)

Transfer learning to object detection and semantic segmentation.

COCO det

Method	Backbone	epoch	AP_bb	AP_50	AP_75	pretrained model	det/seg model
CARE	ResNet50	200	39.4	59.2	42.6	[pretrained] (wip)	[model] (wip)
CARE	ResNet50	400	39.6	59.4	42.9	[pretrained] (wip)	[model] (wip)
CARE	ResNet50-FPN	200	39.5	60.2	43.1	[pretrained] (wip)	[model] (wip)
CARE	ResNet50-FPN	400	39.8	60.5	43.5	[pretrained] (wip)	[model] (wip)

COCO instance seg

Method	Backbone	epoch	AP_mk	AP_50	AP_75	pretrained model	det/seg model
CARE	ResNet50	200	34.6	56.1	36.8	[pretrained] (wip)	[model] (wip)
CARE	ResNet50	400	34.7	56.1	36.9	[pretrained] (wip)	[model] (wip)
CARE	ResNet50-FPN	200	35.9	57.2	38.5	[pretrained] (wip)	[model] (wip)
CARE	ResNet50-FPN	400	36.2	57.4	38.8	[pretrained] (wip)	[model] (wip)

VOC07+12 det

Method	Backbone	epoch	AP_bb	AP_50	AP_75	pretrained model	det/seg model
CARE	ResNet50	200	57.7	83.0	64.5	[pretrained] (wip)	[model] (wip)
CARE	ResNet50	400	57.9	83.0	64.7	[pretrained] (wip)	[model] (wip)

📋 More results are provided in the paper.

Contributing

📋 WIP

Comments

the loss function in paper

Hi, I have read ur paper again today and read with codes, still got a question that why Lc works, our purpose seems to minimize the dissimilarity but the function have a '-' , I think with the '-' it just do the opposite work.

opened by wcyjerry 7

Graph Self-Attention Network for Learning Spatial-Temporal Interaction Representation in Autonomous Driving

GSAN Introduction Code for paper GSAN: Graph Self-Attention Network for Learning Spatial-Temporal Interaction Representation in Autonomous Driving, wh

6 Oct 27, 2022

Pytorch implementation for our ICCV 2021 paper "TRAR: Routing the Attention Spans in Transformers for Visual Question Answering".

TRAnsformer Routing Networks (TRAR) This is an official implementation for ICCV 2021 paper "TRAR: Routing the Attention Spans in Transformers for Visu

49 Nov 10, 2022

This is the official pytorch implementation for our ICCV 2021 paper "TRAR: Routing the Attention Spans in Transformers for Visual Question Answering" on VQA Task

🌈 ERASOR (RA-L'21 with ICRA Option) Official page of "ERASOR: Egocentric Ratio of Pseudo Occupancy-based Dynamic Object Removal for Static 3D Point C

225 Dec 29, 2022

Official PyTorch implementation for paper Context Matters: Graph-based Self-supervised Representation Learning for Medical Images

Context Matters: Graph-based Self-supervised Representation Learning for Medical Images Official PyTorch implementation for paper Context Matters: Gra

49 Nov 23, 2022

[CVPR2021] The source code for our paper 《Removing the Background by Adding the Background: Towards Background Robust Self-supervised Video Representation Learning》.

TBE The source code for our paper "Removing the Background by Adding the Background: Towards Background Robust Self-supervised Video Representation Le

150 Dec 28, 2022

BYOL for Audio: Self-Supervised Learning for General-Purpose Audio Representation

BYOL for Audio: Self-Supervised Learning for General-Purpose Audio Representation This is a demo implementation of BYOL for Audio (BYOL-A), a self-sup

160 Jan 4, 2023

Implementation of Self-supervised Graph-level Representation Learning with Local and Global Structure (ICML 2021).

Self-supervised Graph-level Representation Learning with Local and Global Structure Introduction This project is an implementation of ``Self-supervise

50 Dec 9, 2022

A PyTorch implementation of "Multi-Scale Contrastive Siamese Networks for Self-Supervised Graph Representation Learning", IJCAI-21

MERIT A PyTorch implementation of our IJCAI-21 paper Multi-Scale Contrastive Siamese Networks for Self-Supervised Graph Representation Learning. Depen

Graph Analysis & Deep Learning Laboratory, GRAND

32 Jan 2, 2023

Code for the paper "Spatio-temporal Self-Supervised Representation Learning for 3D Point Clouds" (ICCV 2021)

Spatio-temporal Self-Supervised Representation Learning for 3D Point Clouds This is the official code implementation for the paper "Spatio-temporal Se

63 Jan 5, 2023

Revitalizing CNN Attention via Transformers in Self-Supervised Visual Representation Learning

Related tags

Overview

Revitalizing CNN Attention via Transformers in Self-Supervised Visual Representation Learning

Updates

Requirements

Data Preparation

Training

Evaluation

Pre-trained Models

Model Zoo

Self-supervised learning on image classifications.

Transfer learning to object detection and semantic segmentation.

COCO det

COCO instance seg

VOC07+12 det

Contributing

You might also like...

Graph Self-Attention Network for Learning Spatial-Temporal Interaction Representation in Autonomous Driving

Pytorch implementation for our ICCV 2021 paper "TRAR: Routing the Attention Spans in Transformers for Visual Question Answering".

This is the official pytorch implementation for our ICCV 2021 paper "TRAR: Routing the Attention Spans in Transformers for Visual Question Answering" on VQA Task

Official PyTorch implementation for paper Context Matters: Graph-based Self-supervised Representation Learning for Medical Images

[CVPR2021] The source code for our paper 《Removing the Background by Adding the Background: Towards Background Robust Self-supervised Video Representation Learning》.

BYOL for Audio: Self-Supervised Learning for General-Purpose Audio Representation

Implementation of Self-supervised Graph-level Representation Learning with Local and Global Structure (ICML 2021).

A PyTorch implementation of "Multi-Scale Contrastive Siamese Networks for Self-Supervised Graph Representation Learning", IJCAI-21

Code for the paper "Spatio-temporal Self-Supervised Representation Learning for 3D Point Clouds" (ICCV 2021)

Comments

the loss function in paper

Owner

ChongjianGE

UniMoCo: Unsupervised, Semi-Supervised and Full-Supervised Visual Representation Learning

Implementation of the 😇 Attention layer from the paper, Scaling Local Self-Attention For Parameter Efficient Visual Backbones

An official PyTorch implementation of the TKDE paper "Self-Supervised Graph Representation Learning via Topology Transformations".

Locally Enhanced Self-Attention: Rethinking Self-Attention as Local and Context Terms

The Self-Supervised Learner can be used to train a classifier with fewer labeled examples needed using self-supervised learning.

Weak-supervised Visual Geo-localization via Attention-based Knowledge Distillation

This repository is the official implementation of Unleashing the Power of Contrastive Self-Supervised Visual Models via Contrast-Regularized Fine-Tuning (NeurIPS21).

pytorch implementation of "Contrastive Multiview Coding", "Momentum Contrast for Unsupervised Visual Representation Learning", and "Unsupervised Feature Learning via Non-Parametric Instance-level Discrimination"

Dense Contrastive Learning (DenseCL) for self-supervised representation learning, CVPR 2021.

Self-supervised learning on Graph Representation Learning (node-level task)