OCRA (Object-Centric Recurrent Attention) source code

Hossein Adeli

Last update: Jun 18, 2022

Related tags

Deep Learning OCRA

Overview

OCRA (Object-Centric Recurrent Attention) source code

Hossein Adeli and Seoyoung Ahn

Please cite this article if you find this repository useful:

For data generation and loading
1. stimuli_util.ipynb includes all the codes and the instructions for how to generate the datasets for the three tasks; MultiMNIST, MultiMNIST Cluttered and MultiSVHN.
2. loaddata.py should be updated with the location of the data files for the tasks if not the default used.
For training and testing the model:
1. OCRA_demo.ipynb includes the code for building and training the model. In the first notebook cell, a hyperparameter file should be specified. Parameter files are provided here (different settings are discussed in the supplementary file)
2. multimnist_params_10glimpse.txt and multimnist_params_3glimpse.txt set all the hyperparameters for MultiMNIST task with 10 and 3 glimpses, respectively.
OCRA_demo-MultiMNIST_3glimpse_training.ipynb shows how to load a parameter file and train the model.
1. multimnist_cluttered_params_7glimpse.txt and multimnist_cluttered_params_5glimpse.txt set all the hyperparameters for MultiMNIST Cluttered task with 7 and 5 glimpses, respectively.
2. multisvhn_params.txt sets all the hyperparameters for the MultiSVHN task with 12 glimpses.
3. This notebook also includes code for testing a trained model and also for plotting the attention windows for sample images.
OCRA_demo-cluttered_5steps_loadtrained.ipynb shows how to load a trained model and test it on the test dataset. Example pretrained models are included in the repository under pretrained folder. Download all the pretrained models.

Image-level accuracy averaged from 5 runs

Task (Model name)	Error Rate (SD)
MultiMNIST (OCRA-10glimpse)	5.08 (0.17)
Cluttered MultiMNIST (OCRA-7glimpse)	7.12 (1.05)
MultiSVHN (OCRA-12glimpse)	10.07 (0.53)

Validation losses during training

From MultiMNIST OCRA-10glimpse:

From Cluttered MultiMNIST OCRA-7glimpse

Supplementary Results:

Object-centric behavior
- MultiMNIST Cluttered task with 5 glimpses
- MultiMNIST Cluttered task with 3 glimpses
The Street View House Numbers Dataset

Object-centric behavior

The opportunity to observe the object-centric behavior is bigger in the cluttered task. Since the ratio of the glimpse size to the image size is small (covering less than 4 percent of the image), the model needs to optimally move and select the objects to accurately recognize them. Also reducing the number of glimpses has a similar effect, (we experimented with 3 and 5) forcing the model to leverage its object-centric representation to find the objects without being distracted by the noise segments. We include many more examples of the model behavior with both 3 and 5 glimpses to show this behavior.

MultiMNIST Cluttered task with 5 glimpses

MultiMNIST Cluttered task with 3 glimpses

The Street View House Numbers Dataset

We train the model to "read" the digits from left to right by having the order of the predicted sequence match the ground truth from left to right. We allow the model to make 12 glimpses, with the first two not being constrained and the capsule length from every following two glimpses will be read out for the output digit (e.g. the capsule lengths from the 3rd and 4th glimpses are read out to predict digit number 1; the left-most digit and so on). Below are sample behaviors from our model.

The top five rows show the original images, and the bottom five rows show the reconstructions

The generation of sample images across 12 glimpses

The generatin in a gif fromat

The model learns to detect and reconstruct objects. The model achieved ~2.5 percent error rate on recognizing individual digits and ~10 percent error in recognizing whole sequences still lagging SOTA performance on this measure. We believe this to be strongly related to our small two-layer convolutional backbone and we expect to get better results with a deeper one, which we plan to explore next. However, the model shows reasonable attention behavior in performing this task.

Below shows the model's read and write attention behavior as it reads and reconstructs one image.

Herea are a few sample mistakes from our model:

ground truth [ 1, 10, 10, 10, 10]
prediction [ 0, 10, 10, 10, 10]

ground truth [ 2, 8, 10, 10, 10]
prediction [ 2, 9, 10, 10, 10]

ground truth [ 1, 2, 9, 10, 10]
prediction [ 1, 10, 10, 10, 10]

ground truth [ 5, 1, 10, 10, 10]
prediction [ 5, 7, 10, 10, 10]

Some MNIST cluttered results

Testing the model on MNIST cluttered dataset with three time steps

Code references:

You might also like...

StyleGAN-Human: A Data-Centric Odyssey of Human Generation

StyleGAN-Human: A Data-Centric Odyssey of Human Generation Abstract: Unconditional human image generation is an important task in vision and graphics,

762 Jan 8, 2023

[CVPR 2022 Oral] Versatile Multi-Modal Pre-Training for Human-Centric Perception

Versatile Multi-Modal Pre-Training for Human-Centric Perception Fangzhou Hong1 Liang Pan1 Zhongang Cai1,2,3 Ziwei Liu1* 1S-Lab, Nanyang Technologic

96 Jan 3, 2023

Tools to create pixel-wise object masks, bounding box labels (2D and 3D) and 3D object model (PLY triangle mesh) for object sequences filmed with an RGB-D camera.

Tools to create pixel-wise object masks, bounding box labels (2D and 3D) and 3D object model (PLY triangle mesh) for object sequences filmed with an RGB-D camera. This project prepares training and testing data for various deep learning projects such as 6D object pose estimation projects singleshotpose, as well as object detection and instance segmentation projects.

305 Dec 16, 2022

PyTorch code for our paper "Attention in Attention Network for Image Super-Resolution"

Under construction... Attention in Attention Network for Image Super-Resolution (A2N) This repository is an PyTorch implementation of the paper "Atten

71 Dec 30, 2022

Code for the ECCV2020 paper "A Differentiable Recurrent Surface for Asynchronous Event-Based Data"

A Differentiable Recurrent Surface for Asynchronous Event-Based Data Code for the ECCV2020 paper "A Differentiable Recurrent Surface for Asynchronous

21 Oct 5, 2022

Code and datasets for the paper "Combining Events and Frames using Recurrent Asynchronous Multimodal Networks for Monocular Depth Prediction" (RA-L, 2021)

Combining Events and Frames using Recurrent Asynchronous Multimodal Networks for Monocular Depth Prediction This is the code for the paper Combining E

69 Dec 26, 2022

the code of the paper: Recurrent Multi-view Alignment Network for Unsupervised Surface Registration (CVPR 2021)

RMA-Net This repo is the implementation of the paper: Recurrent Multi-view Alignment Network for Unsupervised Surface Registration (CVPR 2021). Paper

205 Nov 9, 2022

(ICCV 2021) Official code of "Dressing in Order: Recurrent Person Image Generation for Pose Transfer, Virtual Try-on and Outfit Editing."

Dressing in Order (DiOr) 👚 [Paper] 👖 [Webpage] 👗 [Running this code] The official implementation of "Dressing in Order: Recurrent Person Image Gene

277 Dec 28, 2022

Code for the paper "Ordered Neurons: Integrating Tree Structures into Recurrent Neural Networks"

ON-LSTM This repository contains the code used for word-level language model and unsupervised parsing experiments in Ordered Neurons: Integrating Tree

572 Nov 21, 2022

Comments

Data

Hi,

Can you please provide a link to sample data. I tried to generate data with the stimuli ipynb but got the following error message.

tensor_train_ims = torch.Tensor(train_ims)/255 # transform to torch tensor

RuntimeError: [enforce fail at ..\c10\core\CPUAllocator.cpp:79] data. DefaultCPUAllocator: not enough memory: you tried to allocate 15552000000 bytes.

The other other ipynb apparently does not run without the sample data.

Thanks!

opened by dd1github 1

OCRA (Object-Centric Recurrent Attention) source code

Related tags

Overview

OCRA (Object-Centric Recurrent Attention) source code

Image-level accuracy averaged from 5 runs

Validation losses during training

Supplementary Results:

Object-centric behavior

MultiMNIST Cluttered task with 5 glimpses

MultiMNIST Cluttered task with 3 glimpses

The Street View House Numbers Dataset

You might also like...

StyleGAN-Human: A Data-Centric Odyssey of Human Generation

[CVPR 2022 Oral] Versatile Multi-Modal Pre-Training for Human-Centric Perception

Tools to create pixel-wise object masks, bounding box labels (2D and 3D) and 3D object model (PLY triangle mesh) for object sequences filmed with an RGB-D camera.

PyTorch code for our paper "Attention in Attention Network for Image Super-Resolution"

Code for the ECCV2020 paper "A Differentiable Recurrent Surface for Asynchronous Event-Based Data"

Code and datasets for the paper "Combining Events and Frames using Recurrent Asynchronous Multimodal Networks for Monocular Depth Prediction" (RA-L, 2021)

the code of the paper: Recurrent Multi-view Alignment Network for Unsupervised Surface Registration (CVPR 2021)

(ICCV 2021) Official code of "Dressing in Order: Recurrent Person Image Generation for Pose Transfer, Virtual Try-on and Outfit Editing."

Code for the paper "Ordered Neurons: Integrating Tree Structures into Recurrent Neural Networks"

Comments

Data

Owner

Hossein Adeli

The official repo for OC-SORT: Observation-Centric SORT on video Multi-Object Tracking. OC-SORT is simple, online and robust to occlusion/non-linear motion.

Pytorch implementation of "Attention-Based Recurrent Neural Network Models for Joint Intent Detection and Slot Filling"

PyTorch implementation of Hierarchical Multi-label Text Classification: An Attention-based Recurrent Network

Source Code for our paper: Understand me, if you refer to Aspect Knowledge: Knowledge-aware Gated Recurrent Memory Network

Recurrent Scale Approximation (RSA) for Object Detection

NUANCED is a user-centric conversational recommendation dataset that contains 5.1k annotated dialogues and 26k high-quality user turns.

EMNLP'2021: Simple Entity-centric Questions Challenge Dense Retrievers

EMNLP'2021: Simple Entity-centric Questions Challenge Dense Retrievers

Does MAML Only Work via Feature Re-use? A Data Set Centric Perspective

Team nan solution repository for FPT data-centric competition. Data augmentation, Albumentation, Mosaic, Visualization, KNN application