Multi-modal Text Recognition Networks: Interactive Enhancements between Visual and Semantic Features

Last update: Dec 28, 2022

Related tags

Deep Learning MATRN

Overview

Multi-modal Text Recognition Networks: Interactive Enhancements between Visual and Semantic Features

Official PyTorch implementation for Multi-modal Text Recognition Networks: Interactive Enhancements between Visual and Semantic Features (MATRN).

This paper introduces a novel method, called Multi-modAl Text Recognition Network (MATRN), that enables interactions between visual and semantic features for better recognition performances.

Datasets

We use lmdb dataset for training and evaluation dataset. The datasets can be downloaded in clova (for validation and evaluation) and ABINet (for training and evaluation).

Training datasets
Validation datasets
- The union of the training set of ICDAR2013, ICDAR2015, IIIT5K, and Street View Text
Evaluation datasets
- Regular datasets
  - IIIT5K (IIIT)
  - Street View Text (SVT)
  - ICDAR2013: IC13_S with 857 images, IC13_L with 1015 images
- Irregular dataset
  - ICDAR2015: IC15_S with 1811 images, IC15_L with 2077 images
  - Street View Text Perspective (SVTP)
  - CUTE80 (CUTE)

Tree structure of data directory

data
├── charset_36.txt
├── evaluation
│   ├── CUTE80
│   ├── IC13_857
│   ├── IC13_1015
│   ├── IC15_1811
│   ├── IC15_2077
│   ├── IIIT5k_3000
│   ├── SVT
│   └── SVTP
├── training
│   ├── MJ
│   │   ├── MJ_test
│   │   ├── MJ_train
│   │   └── MJ_valid
│   └── ST
├── validation
├── WikiText-103.csv
└── WikiText-103_eval_d1.csv

Requirements

pip install torch==1.7.1 torchvision==0.8.2 fastai==1.0.60 lmdb pillow opencv-python

Pretrained Models

Download pretrained model of MATRN from this link. Performances of the pretrained models are:

Model	IIIT	SVT	IC13_S	IC13_L	IC15_S	IC15_L	SVTP	CUTE
MATRN	96.7	94.9	97.9	95.8	86.6	82.9	90.5	94.1

If you want to train with pretrained visioan and language model, download pretrained model of vision and language model from ABINet (for training and evaluation).

Training and Evaluation

Training

python main.py --config=configs/train_matrn.yaml

Evaluation

python main.py --config=configs/train_matrn.yaml --phase test --image_only

Additional flags:

--checkpoint /path/to/checkpoint set the path of evaluation model
--test_root /path/to/dataset set the path of evaluation dataset
--model_eval [alignment|vision|language] which sub-model to evaluate
--image_only disable dumping visualization of attention masks

Acknowledgements

This implementation has been based on ABINet.

Citation

Please cite this work in your publications if it helps your research.

@article{na2021multi,
  title={Multi-modal Text Recognition Networks: Interactive Enhancements between Visual and Semantic Features},
  author={Na, Byeonghu and Kim, Yoonsik and Park, Sungrae},
  journal={arXiv preprint arXiv:2111.15263},
  year={2021}
}

Comments

Predict More Characters
Hello there!

Great work. I'd like to ask how to train the align model with more characters. The current implementation can only recognize 36 characters (0~9, a~z). I want to recognize 90 characters (0~9, a~z, A~Z, and some symbols).

I tried to modify some code and now I can train on 90 characters. However, I am facing a problem that I can not load the pre-trained language model and vision model, as they are trained on 36 characters. Is there any way to modify the code so that I can load the pre-trained weights?
opened by Mountchicken 4
Question about the performance of pre-trained model that the link contains.

First, thank you very much for your work, it is very impressive. However, when I evaluate with the pre-trained model provided by the link, I get results that are lower than the performance of the report. What is the reason for this? Thank you very much for your answer. And my results on 6 datasets of IIIT5k_3000, SVT, SVTP, IC13_857, IC15_1811, CUTE80 are as follows:

[2022-03-03 23:28:19,374 main.py:276 INFO train-matrn] validation time: 62.44528245925903 / batch size: 384 [2022-03-03 23:28:19,374 main.py:281 INFO train-matrn] eval loss = 1.435, ccr = 0.957, cwr = 0.904, ted = 1542.000, ned = 297, ted/w = 0.213.

you results of the same six datasets average cwr is 93.450.

Thank you very much again!

opened by Zhou2019 4
Question about the usage of text input

I noticed that texts(index encoding) is passed to the forward function, but not used anywhere. Just curious that, are you going to "ADD" text embedding to the final output? My guess is that, you probably have tried it out, but got limited performance improvements. I've been thinking about the usage of text embedding for a while, but it's too hard to convince myself to add text embedding to the training pipeline. As in the inference time, no text information will be given. Please correct me if my guess is wrong. Thanks.

opened by laoShuaiGe 3
Question about reproducing.

Thanks for your great work.

Can you tell me how long the model needs to be trained under the configuration of 4 NVIDIA GeForce RTX 3090GPUs to converge to the results in the paper? If it is convenient, could you provide your training logs?

I'm in the process of reproducing it now, but I found that the loss became jittery after a period of training, I don't know if I configured it wrong or if it's inherently so, slowly converging to the result in a long period of jitters. So I hope the author will provide a training log, if possible（thanks a lot).

Thank you very much!

opened by mrazhou 2
About Code in Line 81 of main.py

Hi! Thanks for your great work. When I debug your code, I find Line81 and Line 82 in file main.py seems like they're written backwards. I have no idea whether this is correct. Thanks :D https://github.com/byeonghu-na/MATRN/blob/f4d43a92555c93df67dbb8c597483e9a5c3fed14/main.py#L81 https://github.com/byeonghu-na/MATRN/blob/f4d43a92555c93df67dbb8c597483e9a5c3fed14/main.py#L82

opened by Gmbition 1
Error about the model when using resnet as the backbone.
Hello author, the following error occurred in the model when I used Resnet as the model backbone instead of ResTransformer. No such error occurred when I ran ABINet using ResNet as the backbone of the model.

RuntimeError: one of the variables needed for gradient computation has been modified by an inplace operation: [torch.FloatTensor [10, 256, 512]], which is output 0 of ViewBackward, is at version 28; expected version 0 instead. Hint: the backtrace further above shows the operation that failed to compute its gradient. The variable in question was changed in there or anywhere later. Good luck!

I found that backbone must have transformer in your model, if there is no transformer behind CNN, there will be mistakes, but I can't find more specific reasons. I only changed the backbone in the train_matrn.yaml configuration file.

Thanks for your reply!
opened by Zhou2019 1
Question about code

Thank you for sharing the code！

The 34-th line in modules/model_matrn_iter.py: the self.semantic_visual has no attribute about pe. So I get the error "torch.nn.modules.module.ModuleAttributeError: 'BaseSemanticVisual_backbone_feature' object has no attribute 'pe'"

So as the 39-th and 44-th lines.

Is there something wrong here？

opened by Sisi0518 1

Multi-modal Text Recognition Networks: Interactive Enhancements between Visual and Semantic Features

Related tags

Overview

Multi-modal Text Recognition Networks: Interactive Enhancements between Visual and Semantic Features

Datasets

Requirements

Pretrained Models

Training and Evaluation

Acknowledgements

Citation

Comments

Predict More Characters

Question about the performance of pre-trained model that the link contains.

Question about the usage of text input

Question about reproducing.

About Code in Line 81 of main.py

Error about the model when using resnet as the backbone.

Question about code

Owner

improvement of CLIP features over the traditional resnet features on the visual question answering, image captioning, navigation and visual entailment tasks.

Pytorch re-implementation of Paper: SwinTextSpotter: Scene Text Spotting via Better Synergy between Text Detection and Text Recognition (CVPR 2022)

Code of U2Fusion: a unified unsupervised image fusion network for multiple image fusion tasks, including multi-modal, multi-exposure and multi-focus image fusion.

Siamese-nn-semantic-text-similarity - A repository containing comprehensive Neural Networks based PyTorch implementations for the semantic text similarity task

AdaMML: Adaptive Multi-Modal Learning for Efficient Video Recognition

Code and pre-trained models for MultiMAE: Multi-modal Multi-task Masked Autoencoders

This is the unofficial code of Deep Dual-resolution Networks for Real-time and Accurate Semantic Segmentation of Road Scenes. which achieve state-of-the-art trade-off between accuracy and speed on cityscapes and camvid, without using inference acceleration and extra data

Conformer: Local Features Coupling Global Representations for Visual Recognition

ROSITA: Enhancing Vision-and-Language Semantic Alignments via Cross- and Intra-modal Knowledge Integration

Code for the paper "MASTER: Multi-Aspect Non-local Network for Scene Text Recognition" (Pattern Recognition 2021)

Image transformations designed for Scene Text Recognition (STR) data augmentation. Published at ICCV 2021 Workshop on Interactive Labeling and Data Augmentation for Vision.

Deep RGB-D Saliency Detection with Depth-Sensitive Attention and Automatic Multi-Modal Fusion (CVPR'2021, Oral)

We present a framework for training multi-modal deep learning models on unlabelled video data by forcing the network to learn invariances to transformations applied to both the audio and video streams.

Multi-modal co-attention for drug-target interaction annotation and Its Application to SARS-CoV-2

Code of paper Interact, Embed, and EnlargE (IEEE): Boosting Modality-specific Representations for Multi-Modal Person Re-identification.

Pytorch Implementation of Zero-Shot Image-to-Text Generation for Visual-Semantic Arithmetic

[CVPR'21] Multi-Modal Fusion Transformer for End-to-End Autonomous Driving

A Multi-modal Model Chinese Spell Checker Released on ACL2021.

A pytorch-based deep learning framework for multi-modal 2D/3D medical image segmentation