Learning Tracking Representations via Dual-Branch Fully Transformer Networks

phiphi

Last update: May 4, 2022

Related tags

Deep Learning DualTFR

Overview

Learning Tracking Representations via Dual-Branch Fully Transformer Networks

DualTFR

⭐ We achieves the runner-ups for both VOT2021ST (short-term) and RT(real-time). The variants of DualTFR take 3rd/4th places of VOT2020RT and 4th places of VOT2020ST

For VOT21 challenge model weight download:

We provide the models of Five trackers SAMN, SAMN_DiMP, DualTFR, DualTFRst, DualTFRon here.

Note that the AlphaRefine (https://github.com/MasterBin-IIAU/AlphaRefine) model and SuperDiMP (https://github.com/visionml/pytracking) model are the same with the original author.

Tracker	model quantity	model name
SAMN	1	SAMN.tar
SAMN_DiMP	2	super_dimp.pth.tar, SAMN.tar
DualTFR	2	DualTFR.tar, ar.pth.tar
DualTFRst	2	DualTFRst.tar, ar.pth.tar
DualTFRon	2	DualTFRon.tar, ar.pth.tar

Models can be downloaded from BaiduNetDisk or GoogleDrive:

BaiduNetDisk:

https://pan.baidu.com/s/1RHA7HVlXtNEzYPGIjJbQ-g (sruh)

GoogleDrive:

https://drive.google.com/drive/folders/1Z61_mfh2vwzqDxejt5idBOgYhWOCZOr5?usp=sharing

Code will be released soon.

We present a simple Siamese-like Dual-branch network based on solely Transformer networks to learn about tracking features. Given a template and a search image, we divide them into non-overlapping image patches and extract a feature vector for each based on its matching results with others within an attention window. Then for each token, we estimate whether it contains the target object and the corresponding size. The prominent advantage of the approach is that the features are learned from matching, and ultimately, for matching. So the features are aligned with the subsequent object tracking task. The method achieves comparable results comparing to the best-performing methods which first use CNN to extract features and then use Transformer to fuse them. Without bells and whistles, it outperforms the state-of-the-art methods on GOT-10k and VOT2020 benchmarks. In addition, the method achieves real-time inference speed (about 40 fps).

Acknowledgments

Thanks for the great PyTracking Library, which helps us to quickly implement our ideas.
We use the implementation of the Swin Transformer from the official repo https://github.com/microsoft/Swin-Transformer.

Contacts

Fei Xie, School of Automation, Southeast University, China, [email protected], wechat: 372998044

Comments

About train details

Hello, I'd like to ask you for specific training details in this job. You said you used 8 GPU and each epoch used 40 million pairs of samples. Then, did all the epochs add up to 4 billion? How long did you train in total?

opened by wjc0602 1

An official source code for paper Deep Graph Clustering via Dual Correlation Reduction, accepted by AAAI 2022

Dual Correlation Reduction Network An official source code for paper Deep Graph Clustering via Dual Correlation Reduction, accepted by AAAI 2022. Any

109 Dec 23, 2022

This is the unofficial code of Deep Dual-resolution Networks for Real-time and Accurate Semantic Segmentation of Road Scenes. which achieve state-of-the-art trade-off between accuracy and speed on cityscapes and camvid, without using inference acceleration and extra data

Deep Dual-resolution Networks for Real-time and Accurate Semantic Segmentation of Road Scenes Introduction This is the unofficial code of Deep Dual-re

113 Dec 23, 2022

[TIP 2021] SADRNet: Self-Aligned Dual Face Regression Networks for Robust 3D Dense Face Alignment and Reconstruction

Learning Tracking Representations via Dual-Branch Fully Transformer Networks

Related tags

Overview

Learning Tracking Representations via Dual-Branch Fully Transformer Networks

DualTFR

For VOT21 challenge model weight download:

Code will be released soon.

Acknowledgments

Contacts

You might also like...

An official source code for paper Deep Graph Clustering via Dual Correlation Reduction, accepted by AAAI 2022

This is the unofficial code of Deep Dual-resolution Networks for Real-time and Accurate Semantic Segmentation of Road Scenes. which achieve state-of-the-art trade-off between accuracy and speed on cityscapes and camvid, without using inference acceleration and extra data

[TIP 2021] SADRNet: Self-Aligned Dual Face Regression Networks for Robust 3D Dense Face Alignment and Reconstruction

Joint detection and tracking model named DEFT, or ``Detection Embeddings for Tracking.

Tracking code for the winner of track 1 in the MMP-Tracking Challenge at ICCV 2021 Workshop.

Tracking Pipeline helps you to solve the tracking problem more easily

Quadruped-command-tracking-controller - Quadruped command tracking controller (flat terrain)

Python package for multiple object tracking research with focus on laboratory animals tracking.

VSR-Transformer - This paper proposes a new Transformer for video super-resolution (called VSR-Transformer).

Comments

About train details

Owner

phiphi

[SIGGRAPH Asia 2021] DeepVecFont: Synthesizing High-quality Vector Fonts via Dual-modality Learning.

Diverse Branch Block: Building a Convolution as an Inception-like Unit

I decide to sync up this repo and self-critical.pytorch. (The old master is in old master branch for archive)

Angora is a mutation-based fuzzer. The main goal of Angora is to increase branch coverage by solving path constraints without symbolic execution.

Only works with the dashboard version / branch of jesse

The official implementation of paper Siamese Transformer Pyramid Networks for Real-Time UAV Tracking, accepted by WACV22

[CVPR 2022] TransEditor: Transformer-Based Dual-Space GAN for Highly Controllable Facial Editing

Code repository for the paper "Tracking People with 3D Representations"

[ICCV'2021] Image Inpainting via Conditional Texture and Structure Dual Generation