Improving Object Detection by Estimating Bounding Box Quality Accurately

Last update: Apr 14, 2022

Related tags

Deep Learning LQM

Overview

Improving Object Detection by Estimating Bounding Box Quality Accurately

Abstract

Object detection aims to locate and classify object instances in images. Therefore, the object detection model is generally implemented with two parallel branches to optimize localization and classification. After training the detection model, we should select the best bounding box of each class among a number of estimations for reliable inference. Generally, NMS (Non Maximum Suppression) is operated to suppress low-quality bounding boxes by referring to classification scores or center-ness scores. However, since the quality of bounding boxes is not considered, the low-quality bounding boxes can be accidentally selected as a positive bounding box for the corresponding class. We believe that this misalignment between two parallel tasks causes degrading of the object detection performance. In this paper, we propose a method to estimate bounding boxes' quality using four-directional Gaussian quality modeling, which leads the consistent results between two parallel branches. Extensive experiments on the MS COCO benchmark show that the proposed method consistently outperforms the baseline (FCOS). Eventually, our best model offers the state-of-the-art performance by achieving 48.9% in AP. We also confirm the efficiency of the method by comparing the number of parameters and computational overhead.

Overall Architecture

Implementation Details

We implement our detection model on top of MMDetection (v2.6), an open source object detection toolbox. If not specified separately, the default settings of FCOS implementation are not changed. We train and validate our network on four RTX TITAN GPUs in the environment of Pytorch v1.6 and CUDA v10.2.

Please see GETTING_STARTED.md for the basic usage of MMDetection.

Installation

Clone the this repository.

git clone https://github.com/sanghun3819/LQM.git
cd LQM

Create a conda virtural environment and install dependencies.
```
conda env create -f environment.yml
```
Activate conda environment
```
conda activate lqm
```

Install build requirements and then install MMDetection.

pip install -r requirements/build.txt
pip install -v -e .

Preparing MS COCO dataset

bash download_coco.sh

Preparing Pre-trained model weights

bash download_weights.sh

Train

# assume that you are under the root directory of this project,
# and you have activated your virtual environment if needed.
# and with COCO dataset in 'data/coco/'

./tools/dist_train.sh configs/uncertainty_guide/uncertainty_guide_r50_fpn_1x.py 4 --validate

Inference

./tools/dist_test.sh configs/uncertainty_guide/uncertainty_guide_r50_fpn_1x.py work_dirs/uncertainty_guide_r50_fpn_1x/epoch_12.pth 4 --eval bbox

Image demo using pretrained model weight

# Result will be saved under the demo directory of this project (detection_result.jpg)
# config, checkpoint, source image path are needed (If you need pre-trained weights, you can download them from provided google drive link)
# score threshold is optional

python demo/LQM_image_demo.py --config configs/uncertainty_guide/uncertainty_guide_r50_fpn_1x.py --checkpoint work_dirs/pretrained/LQM_r50_fpn_1x.pth --img data/coco/test2017/000000011245.jpg --score-thr 0.3

Models

For your convenience, we provide the following trained models. All models are trained with 16 images in a mini-batch with 4 GPUs.

Model	Multi-scale training	AP (minival)	Link
LQM_R50_FPN_1x	No	40.0	Google
LQM_R101_FPN_2x	Yes	44.8	Google
LQM_R101_dcnv2_FPN_2x	Yes	47.4	Google
LQM_X101_FPN_2x	Yes	47.2	Google
LQM_X101_dcnv2_FPN_2x	Yes	48.9	Google

ByteTrack(Multi-Object Tracking by Associating Every Detection Box)のPythonでのONNX推論サンプル

ByteTrack-ONNX-Sample ByteTrack(Multi-Object Tracking by Associating Every Detection Box)のPythonでのONNX推論サンプルです。 ONNXに変換したモデルも同梱しています。変換自体を試したい方はByteT

16 Oct 26, 2022

Implementation of Analyzing and Improving the Image Quality of StyleGAN (StyleGAN 2) in PyTorch

2.2k Jan 1, 2023

Implementation of Common Image Evaluation Metrics by Sayed Nadim (sayednadim.github.io). The repo is built based on full reference image quality metrics such as L1, L2, PSNR, SSIM, LPIPS. and feature-level quality metrics such as FID, IS. It can be used for evaluating image denoising, colorization, inpainting, deraining, dehazing etc. where we have access to ground truth.

Image Quality Evaluation Metrics Implementation of some common full reference image quality metrics. The repo is built based on full reference image q

10 Jan 1, 2023

Improving 3D Object Detection with Channel-wise Transformer

"Improving 3D Object Detection with Channel-wise Transformer" Thanks for the OpenPCDet, this implementation of the CT3D is mainly based on the pcdet v

107 Dec 20, 2022

Code for the paper SphereRPN: Learning Spheres for High-Quality Region Proposals on 3D Point Clouds Object Detection, ICIP 2021.

SphereRPN Code for the paper SphereRPN: Learning Spheres for High-Quality Region Proposals on 3D Point Clouds Object Detection, ICIP 2021. Authors: Th

15 Dec 2, 2022

Hybrid CenterNet - Hybrid-supervised object detection / Weakly semi-supervised object detection

Improving Object Detection by Estimating Bounding Box Quality Accurately

Related tags

Overview

Improving Object Detection by Estimating Bounding Box Quality Accurately

Abstract

Overall Architecture

Implementation Details

Installation

Preparing MS COCO dataset

Preparing Pre-trained model weights

Train

Inference

Image demo using pretrained model weight

Models

You might also like...

ByteTrack(Multi-Object Tracking by Associating Every Detection Box)のPythonでのONNX推論サンプル

Implementation of Analyzing and Improving the Image Quality of StyleGAN (StyleGAN 2) in PyTorch

Improving 3D Object Detection with Channel-wise Transformer

Code for the paper SphereRPN: Learning Spheres for High-Quality Region Proposals on 3D Point Clouds Object Detection, ICIP 2021.

Hybrid CenterNet - Hybrid-supervised object detection / Weakly semi-supervised object detection

Yolo object detection - Yolo object detection with python

labelpix is a graphical image labeling interface for drawing bounding boxes

Pytorch based library to rank predicted bounding boxes using text/image user's prompts.

Owner

Tools to create pixel-wise object masks, bounding box labels (2D and 3D) and 3D object model (PLY triangle mesh) for object sequences filmed with an RGB-D camera.

This code finds bounding box of a single human mouth.

Alpha-IoU: A Family of Power Intersection over Union Losses for Bounding Box Regression

Fast algorithms to compute an approximation of the minimal volume oriented bounding box of a point cloud in 3D.

Black-Box-Tuning - Black-Box Tuning for Language-Model-as-a-Service

A scikit-learn-compatible module for estimating prediction intervals.

A DNN inference latency prediction toolkit for accurately modeling and predicting the latency on diverse edge devices.

Estimating and Exploiting the Aleatoric Uncertainty in Surface Normal Estimation

Official codes: Self-Supervised Learning by Estimating Twin Class Distribution

This is a Keras implementation of a CNN for estimating age, gender and mask from a camera.