This is the official released code for our paper, The Emergence of Objectness: Learning Zero-Shot Segmentation from Videos

Last update: Oct 8, 2022

Related tags

Deep Learning The-Emergence-of-Objectness

Overview

The-Emergence-of-Objectness

This is the official released code for our paper, The Emergence of Objectness: Learning Zero-Shot Segmentation from Videos, which has been accepted by NeurIPS 2021. Code will be available soon.

Code

To be released.

Abstract

Humans can easily segment moving objects without knowing what they are. That objectness could emerge from continuous visual observations motivates us to model grouping and movement concurrently from unlabeled videos. Our premise is that a video has different views of the same scene related by moving components, and the right region segmentation and region flow would allow mutual view synthesis which can be checked from the data itself without any external supervision.

Our model starts with two separate pathways: an appearance pathway that outputs feature-based region segmentation for a single image, and a motion pathway that outputs motion features for a pair of images. It then binds them in a conjoint representation called segment flow that pools flow offsets over each region and provides a gross characterization of moving regions for the entire scene. By training the model to minimize view synthesis errors based on segment flow, our appearance and motion pathways learn region segmentation and flow estimation automatically without building them up from low-level edges or optical flows respectively.

Our model demonstrates the surprising emergence of objectness in the appearance pathway, surpassing prior works on zero-shot object segmentation from an image, moving object segmentation from a video with unsupervised test-time adaptation, and semantic image segmentation by supervised fine-tuning. Our work is the first truly end-to-end zero-shot object segmentation from videos. It not only develops generic objectness for segmentation and tracking, but also outperforms prevalent image-based contrastive learning methods without augmentation engineering.

Approach

We learn a single-image segmentation network and a dual-frame motion network with an unsupervised image reconstruction loss. We sample two frames, $i$ and $j$, from a video. Frame $i$ goes through the segmentation network and outputs a set of masks, whereas frames $i$ and $j$ go through the motion network and output a feature map. The feature is pooled per mask and a flow is predicted. All the segments and their flows are combined into a segment flow representation from frame $i$ → $j$, which are used to warp frame $i$ into $j$, and compared against frame $j$ to train the two networks.

Zero-Shot Saliency Detection

Qualitative salient object detection results. We directly transfer our pretrained segmentation network to novel images on the DUTS dataset without any finetuning. Surprisingly, we find that the model pretrained on videos to segment moving objects can generalize to detect stationary unmovable objects in a static image, e.g. the statue, the plate, the bench and the tree in the last column.

Zero-shot Video Object Segmentation

Qualitative results of SegTrackv2

Qualitative results of DAVIS 2016

Qualitative results of FBMS59

You might also like...

Zero-shot Synthesis with Group-Supervised Learning (ICLR 2021 paper)

Comments

Question about Segmentation Network

Hello, I appreciate your awesome work. In paper, output channels of segmentation network is 5 channels. What channels is used to predict SOD binary mask?? Can you guide parts of segmentation network in your code??

opened by iseunghoon 0

Dataset

JPEGImages/480p/blackswan/ 00000.jpg 00001.jpg 00002.jpg 00003.jpg 00004.jpg 00005.jpg 
JPEGImages/480p/blackswan/ 00000.jpg 00001.jpg 00002.jpg 00003.jpg 00004.jpg 00005.jpg 
JPEGImages/480p/blackswan/ 00000.jpg 00001.jpg 00002.jpg 00003.jpg 00004.jpg 00005.jpg 
JPEGImages/480p/blackswan/ 00000.jpg 00001.jpg 00002.jpg 00003.jpg 00004.jpg 00005.jpg 
JPEGImages/480p/blackswan/ 00000.jpg 00001.jpg 00002.jpg 00003.jpg 00004.jpg 00005.jpg 
JPEGImages/480p/blackswan/ 00000.jpg 00001.jpg 00002.jpg 00003.jpg 00004.jpg 00005.jpg 
JPEGImages/480p/blackswan/ 00000.jpg 00001.jpg 00002.jpg 00003.jpg 00004.jpg 00005.jpg 
JPEGImages/480p/blackswan/ 00000.jpg 00001.jpg 00002.jpg 00003.jpg 00004.jpg 00005.jpg 
JPEGImages/480p/blackswan/ 00000.jpg 00001.jpg 00002.jpg 00003.jpg 00004.jpg 00005.jpg 
JPEGImages/480p/blackswan/ 00000.jpg 00001.jpg 00002.jpg 00003.jpg 00004.jpg 00005.jpg

I want to konw why are so many lines repeated, thanks

opened by GuoQuanhao 0

Open source time

Hello, author, I am very excited for you to do such a valuable work. Is there any specific arrangement for the open source time of the code? Because as far as I know, the NIPS meeting has already been held, and I look forward to seeing your code, thank you!

opened by fupiao1998 0

This is the official released code for our paper, The Emergence of Objectness: Learning Zero-Shot Segmentation from Videos

Related tags

Overview

The-Emergence-of-Objectness

Code

Abstract

Approach

Zero-Shot Saliency Detection

Zero-shot Video Object Segmentation

Qualitative results of SegTrackv2

Qualitative results of DAVIS 2016

Qualitative results of FBMS59

You might also like...

Zero-shot Synthesis with Group-Supervised Learning (ICLR 2021 paper)

NP DRAW paper released code

The code of Zero-shot learning for low-light image enhancement based on dual iteration

Codes for ACL-IJCNLP 2021 Paper "Zero-shot Fact Verification by Claim Generation"

Original code for "Zero-Shot Domain Adaptation with a Physics Prior"

The source code for Generating Training Data with Language Models: Towards Zero-Shot Language Understanding.

Shared Attention for Multi-label Zero-shot Learning

PyTorch implementation of 1712.06087 "Zero-Shot" Super-Resolution using Deep Internal Learning

EMNLP 2021 Adapting Language Models for Zero-shot Learning by Meta-tuning on Dataset and Prompt Collections

Comments

Question about Segmentation Network

Dataset

Open source time

Owner

code for CVPR paper Zero-shot Instance Segmentation

Public repository of the 3DV 2021 paper "Generative Zero-Shot Learning for Semantic Segmentation of 3D Point Clouds"

An official implementation of "Exploiting a Joint Embedding Space for Generalized Zero-Shot Semantic Segmentation" (ICCV 2021) in PyTorch.

Official Pytorch Implementation of: "Semantic Diversity Learning for Zero-Shot Multi-label Classification"(2021) paper

Code for our method RePRI for Few-Shot Segmentation. Paper at http://arxiv.org/abs/2012.06166

[ICCV 2021] Official Pytorch implementation for Discriminative Region-based Multi-Label Zero-Shot Learning SOTA results on NUS-WIDE and OpenImages

[ICCV 2021] Official Pytorch implementation for Discriminative Region-based Multi-Label Zero-Shot Learning SOTA results on NUS-WIDE and OpenImages

Zsseg.baseline - Zero-Shot Semantic Segmentation

Code repo for EMNLP21 paper "Zero-Shot Information Extraction as a Unified Text-to-Triple Translation"

Code for the AAAI 2022 paper "Zero-Shot Cross-Lingual Machine Reading Comprehension via Inter-Sentence Dependency Graph".