Code for "Unsupervised Layered Image Decomposition into Object Prototypes" paper

Last update: Dec 22, 2022

Related tags

Deep Learning computer-vision deep-learning unsupervised sprites unsupervised-learning image-decomposition multi-object

Overview

DTI-Sprites

Pytorch implementation of "Unsupervised Layered Image Decomposition into Object Prototypes" paper

Check out our paper and webpage for details!

If you find this code useful in your research, please cite:

@article{monnier2021dtisprites,
  title={{Unsupervised Layered Image Decomposition into Object Prototypes}},
  author={Monnier, Tom and Vincent, Elliot and Ponce, Jean and Aubry, Mathieu},
  journal={arXiv},
  year={2021},
}

Installation 👷

1. Create conda environment

conda env create -f environment.yml
conda activate dti-sprites

Optional: some monitoring routines are implemented, you can use them by specifying the visdom port in the config file. You will need to install visdom from source beforehand

git clone https://github.com/facebookresearch/visdom
cd visdom && pip install -e .

2. Download non-torchvision datasets

./download_data.sh

This command will download following datasets:

Tetrominoes, Multi-dSprites and CLEVR6 (link to the original repo multi-object datasets with raw tfrecords)
GTSRB (link to the original dataset page)
Weizmann Horse database (link to the original dataset page)
Instagram collections associated to #santaphoto and #weddingkiss (link to the original repo with datasets links and descriptions)

NB: it may happen that gdown hangs, if so you can download them by hand with following gdrive links, unzip and move them to the datasets folder:

How to use 🚀

1. Launch a training

cuda=gpu_id config=filename.yml tag=run_tag ./pipeline.sh

where:

gpu_id is a target cuda device id,
filename.yml is a YAML config located in configs folder,
run_tag is a tag for the experiment.

Results are saved at runs/${DATASET}/${DATE}_${run_tag} where DATASET is the dataset name specified in filename.yml and DATE is the current date in mmdd format. Some training visual results like sprites evolution and reconstruction examples will be saved. Here is an example from Tetrominoes dataset:

Reconstruction examples

Sprites evolution and final

More visual results are available at https://imagine.enpc.fr/~monniert/DTI-Sprites/extra_results/.

2. Reproduce our quantitative results

To launch 5 runs on Tetrominoes benchmark and reproduce our results:

cuda=gpu_id config=tetro.yml tag=default ./multi_pipeline.sh

Available configs are:

Multi-object benchmarks: tetro.yml, dpsrites_gray.yml, clevr6.yml
Clustering benchmarks: gtsrb8.yml, svhn.yml
Cosegmentation dataset: horse.yml

3. Reproduce our qualitative results on Instagram collections

(skip if already downloaded with script above) Create a santaphoto dataset by running process_insta_santa.sh script. It can take a while to scrape the 10k posts from Instagram.
Launch training with cuda=gpu_id config=instagram.yml tag=santaphoto ./pipeline.sh

That's it!

Top 8 sprites discovered

Decomposition examples

Further information

If you like this project, please check out related works on deep transformations from our group:

Code for the paper Learning the Predictability of the Future

Learning the Predictability of the Future Code from the paper Learning the Predictability of the Future. Website of the project in hyperfuture.cs.colu

Computer Vision Lab at Columbia University

139 Nov 18, 2022

PyTorch code for the paper: FeatMatch: Feature-Based Augmentation for Semi-Supervised Learning

FeatMatch: Feature-Based Augmentation for Semi-Supervised Learning This is the PyTorch implementation of our paper: FeatMatch: Feature-Based Augmentat

43 Nov 19, 2022

Code for the paper A Theoretical Analysis of the Repetition Problem in Text Generation

A Theoretical Analysis of the Repetition Problem in Text Generation This repository share the code for the paper "A Theoretical Analysis of the Repeti

37 Nov 21, 2022

Code for our ICASSP 2021 paper: SA-Net: Shuffle Attention for Deep Convolutional Neural Networks

SA-Net: Shuffle Attention for Deep Convolutional Neural Networks (paper) By Qing-Long Zhang and Yu-Bin Yang [State Key Laboratory for Novel Software T

199 Jan 8, 2023

Open source repository for the code accompanying the paper 'Non-Rigid Neural Radiance Fields Reconstruction and Novel View Synthesis of a Deforming Scene from Monocular Video'.

Non-Rigid Neural Radiance Fields This is the official repository for the project "Non-Rigid Neural Radiance Fields: Reconstruction and Novel View Synt

296 Dec 29, 2022

Code for the Shortformer model, from the paper by Ofir Press, Noah A. Smith and Mike Lewis.

Shortformer This repository contains the code and the final checkpoint of the Shortformer model. This file explains how to run our experiments on the

138 Apr 15, 2022

PyTorch code for ICLR 2021 paper Unbiased Teacher for Semi-Supervised Object Detection

Unbiased Teacher for Semi-Supervised Object Detection This is the PyTorch implementation of our paper: Unbiased Teacher for Semi-Supervised Object Detection

366 Dec 28, 2022

Official code for paper "Optimization for Oriented Object Detection via Representation Invariance Loss".

Optimization for Oriented Object Detection via Representation Invariance Loss By Qi Ming, Zhiqiang Zhou, Lingjuan Miao, Xue Yang, and Yunpeng Dong. Th

56 Nov 28, 2022

Code for our CVPR 2021 paper "MetaCam+DSCE"

Joint Noise-Tolerant Learning and Meta Camera Shift Adaptation for Unsupervised Person Re-Identification (CVPR'21) Introduction Code for our CVPR 2021

59 Oct 31, 2022

Comments

morphological transformation for RGB sprites?

Hi -- First, thanks for putting up the code and material here. It's very interesting work and this code repo has been great of help to understand your model.

While I'm experimenting with your codes, I noticed that the current implementation of morphological transformation does not work with color sprites. It gives a size mismatch in the first dimension of sprites and alpha/kernel tensors which have B instead of B*C that sprites have. Just wonder if there's any workaround this.

Also, can you provide little more details on how the dataset is generated? especially the mask labels? I'm trying to run the model on overlapping mnist dataset (as shown in your pipeline figure, but it give me a warning like below: WARN InstanceSegScores._fast_hist error: labels in GT are greater than nb instances

Any input will be appreciated! Thank you very much.

opened by ahnchive 2

RuntimeError: multiple tensor pointing at the same memory location

Hey there! I loved your paper and I'm finding your code quite interesting.

I'm running a training job locally on the Tetrominoes config and dataset and, while first iterations run smoothly, at the end of the epoch I'm finding this error, which seems to stem from the optimization.step() line.

Traceback (most recent call last):
  File "src/trainer.py", line 1091, in <module>
    trainer.run(seed=seed)
  File "/home/ubuntu/education/src/dti-sprites/src/utils/__init__.py", line 102, in wrapper
    return f(*args, **kw)
  File "src/trainer.py", line 292, in run
    self.single_train_batch_run(images)
  File "src/trainer.py", line 337, in single_train_batch_run
    self.optimizer.step()
  File "/home/ubuntu/anaconda3/envs/dti-sprites/lib/python3.7/site-packages/torch/optim/lr_scheduler.py", line 67, in wrapper
    return wrapped(*args, **kwargs)
  File "/home/ubuntu/anaconda3/envs/dti-sprites/lib/python3.7/site-packages/torch/autograd/grad_mode.py", line 15, in decorate_context
    return func(*args, **kwargs)
  File "/home/ubuntu/anaconda3/envs/dti-sprites/lib/python3.7/site-packages/torch/optim/adam.py", line 131, in step
    p.addcdiv_(exp_avg, denom, value=-step_size)
RuntimeError: unsupported operation: more than one element of the written-to tensor refers to a single memory location. 
Please clone() the tensor before performing the operation.

Have you had this same issue? Do you know how to solve it? I believe it's caused by a tensor not being called the clone() or expand() operation.

opened by JHevia23 2

src/model/dti_sprites.py, there is no self.occ_noise

Hello! When I debug the trainer.py with config=dsprites.yml, there was an error

It's about missing self.occ_noise code at all in if self.occ_noise > 0 and self.training: (247 line in model/dti_sprites.py)

If you can find any mistakes, please modify and notice about that. Thanx

opened by actruce 1

Code for "Unsupervised Layered Image Decomposition into Object Prototypes" paper

Related tags

Overview

DTI-Sprites

Installation 👷

1. Create conda environment

2. Download non-torchvision datasets

How to use 🚀

1. Launch a training

Reconstruction examples

Sprites evolution and final

2. Reproduce our quantitative results

3. Reproduce our qualitative results on Instagram collections

Top 8 sprites discovered

Decomposition examples

Further information

You might also like...

Code for the paper Learning the Predictability of the Future

PyTorch code for the paper: FeatMatch: Feature-Based Augmentation for Semi-Supervised Learning

Code for the paper A Theoretical Analysis of the Repetition Problem in Text Generation

Code for our ICASSP 2021 paper: SA-Net: Shuffle Attention for Deep Convolutional Neural Networks

Open source repository for the code accompanying the paper 'Non-Rigid Neural Radiance Fields Reconstruction and Novel View Synthesis of a Deforming Scene from Monocular Video'.

Code for the Shortformer model, from the paper by Ofir Press, Noah A. Smith and Mike Lewis.

PyTorch code for ICLR 2021 paper Unbiased Teacher for Semi-Supervised Object Detection

Official code for paper "Optimization for Oriented Object Detection via Representation Invariance Loss".

Code for our CVPR 2021 paper "MetaCam+DSCE"

Comments

morphological transformation for RGB sprites?

RuntimeError: multiple tensor pointing at the same memory location

src/model/dti_sprites.py, there is no self.occ_noise

Owner

Inference code for "StylePeople: A Generative Model of Fullbody Human Avatars" paper. This code is for the part of the paper describing video-based avatars.

This is the official source code for SLATE. We provide the code for the model, the training code, and a dataset loader for the 3D Shapes dataset. This code is implemented in Pytorch.

Code for paper ECCV 2020 paper: Who Left the Dogs Out? 3D Animal Reconstruction with Expectation Maximization in the Loop.

TensorFlow code for the neural network presented in the paper: "Structural Language Models of Code" (ICML'2020)

Code for the prototype tool in our paper "CoProtector: Protect Open-Source Code against Unauthorized Training Usage with Data Poisoning".

Code to use Augmented Shapiro Wilks Stopping, as well as code for the paper "Statistically Signifigant Stopping of Neural Network Training"

Code for our method RePRI for Few-Shot Segmentation. Paper at http://arxiv.org/abs/2012.06166

Code for ACM MM 2020 paper "NOH-NMS: Improving Pedestrian Detection by Nearby Objects Hallucination"

Official TensorFlow code for the forthcoming paper

This is the code for the paper "Contrastive Clustering" (AAAI 2021)