[LREC] MMChat: Multi-Modal Chat Dataset on Social Media

Silver

Last update: Jan 3, 2023

Related tags

Overview

MMChat

This repo contains the code and data for the LREC2022 paper MMChat: Multi-Modal Chat Dataset on Social Media.

Dataset

MMChat is a large-scale dialogue dataset that contains image-grounded dialogues in Chinese. Each dialogue in MMChat is associated with one or more images (maximum 9 images per dialogue). We design various strategies to ensure the quality of the dialogues in MMChat. Please read our paper for more details. The images in the dataset are hosted on Weibo's static image server. You can refer to the scripts provided in data_processing/weibo_image_crawler to download these images.

Two sample dialogues form MMChat are given below (translated from Chinese):

MMChat is released in different versions:

Rule Filtered Raw MMChat

This version of MMChat contains raw dialogues filtered by our rules. The following table shows some basic statistics:

Item Description	Count
Sessions	4.257 M
Sessions with more than 4 utterances	2.304 M
Utterances	18.590 M
Images	4.874 M
Avg. utterance per session	4.367
Avg. image per session	1.670
Avg. character per utterance	14.104

We devide above dialogues into 9 splits to facilitate the download:

LCCC Filtered MMChat

This version of MMChat contains the dialogues that are filtered based on the LCCC (Large-scale Cleaned Chinese Conversation) dataset. Specifically, some dialogues in MMChat are also contained in LCCC. We regard these dialogues as cleaner dialogues since sophisticated schemes are designed in LCCC to filter out noises. This version of MMChat is obtained using the script data_processing/LCCC_filter.py The following table shows some basic statistics:

Item Description	Count
Sessions	492.6 K
Sessions with more than 4 utterances	208.8 K
Utterances	1.986 M
Images	1.066 M
Avg. utterance per session	4.031
Avg. image per session	2.514
Avg. character per utterance	11.336

We devide above dialogues into 9 splits to facilitate the download:

MMChat

The MMChat dataset reported in our paper are given here. The Weibo content corresponding to these dialogues are all "分享图片", (i.e., "Share Images" in English). The following table shows some basic statistics:

Item Description	Count
Sessions	120.84 K
Sessions with more than 4 utterances	17.32 K
Utterances	314.13 K
Images	198.82 K
Avg. utterance per session	2.599
Avg. image per session	2.791
Avg. character per utterance	8.521

The above dialogues can be downloaded from either Google Drive or Baidu Netdisk.

MMChat-hf

We perform human annotation on the sampled dialogues to determine whether the given images are related to the corresponding dialogues. The following table only shows the statistics for dialogues that are annotated as image-related.

Item Description	Count
Sessions	19.90 K
Sessions with more than 4 utterances	8.91 K
Utterances	81.06 K
Images	52.66K
Avg. utterance per session	4.07
Avg. image per session	2.70
Avg. character per utterance	11.93

We annotated about 100K dialogues. All the annotated dialogues can be downloaded from either Google Drive or Baidu Netdisk.

Code

We are also releasing all the codes used for our experiments. You can use the script run_training.sh in each folder to launch the distributed training.

For models that require image features, you can extract the image features using the scripts in data_processing/extract_image_features

The model shown in our paper can be found in dialog_image:

Reference

Please cite our paper if you find our work useful ;)

@inproceedings{zheng2022MMChat,
  author    = {Zheng, Yinhe and Chen, Guanyi and Liu, Xin and Sun, Jian},
  title     = {MMChat: Multi-Modal Chat Dataset on Social Media},
  booktitle = {Proceedings of The 13th Language Resources and Evaluation Conference},
  year      = {2022},
  publisher = {European Language Resources Association},
}

@inproceedings{wang2020chinese,
  title     = {A Large-Scale Chinese Short-Text Conversation Dataset},
  author    = {Wang, Yida and Ke, Pei and Zheng, Yinhe and Huang, Kaili and Jiang, Yong and Zhu, Xiaoyan and Huang, Minlie},
  booktitle = {NLPCC},
  year      = {2020},
  url       = {https://arxiv.org/abs/2008.03946}
}

We present a framework for training multi-modal deep learning models on unlabelled video data by forcing the network to learn invariances to transformations applied to both the audio and video streams.

Multi-Modal Self-Supervision using GDT and StiCa This is an official pytorch implementation of papers: Multi-modal Self-Supervision from Generalized D

42 Dec 9, 2022

A pytorch-based deep learning framework for multi-modal 2D/3D medical image segmentation

A 3D multi-modal medical image segmentation library in PyTorch We strongly believe in open and reproducible deep learning research. Our goal is to imp

1.2k Dec 27, 2022

AdaMML: Adaptive Multi-Modal Learning for Efficient Video Recognition

AdaMML: Adaptive Multi-Modal Learning for Efficient Video Recognition [ArXiv] [Project Page] This repository is the official implementation of AdaMML:

43 Dec 26, 2022

Self-supervised Multi-modal Hybrid Fusion Network for Brain Tumor Segmentation

JBHI-Pytorch This repository contains a reference implementation of the algorithms described in our paper "Self-supervised Multi-modal Hybrid Fusion N

5 Dec 13, 2021

Multi-modal co-attention for drug-target interaction annotation and Its Application to SARS-CoV-2

CoaDTI Multi-modal co-attention for drug-target interaction annotation and Its Application to SARS-CoV-2 Abstract Environment The test was conducted i

7 Nov 14, 2022

Code of paper Interact, Embed, and EnlargE (IEEE): Boosting Modality-specific Representations for Multi-Modal Person Re-identification.

Interact, Embed, and EnlargE (IEEE): Boosting Modality-specific Representations for Multi-Modal Person Re-identification We provide the codes for repr

12 Dec 12, 2022

Multi-Modal Machine Learning toolkit based on PyTorch.

[LREC] MMChat: Multi-Modal Chat Dataset on Social Media

Related tags

Overview

MMChat

Dataset

Rule Filtered Raw MMChat

LCCC Filtered MMChat

MMChat

MMChat-hf

Code

Reference

You might also like...

We present a framework for training multi-modal deep learning models on unlabelled video data by forcing the network to learn invariances to transformations applied to both the audio and video streams.

A pytorch-based deep learning framework for multi-modal 2D/3D medical image segmentation

AdaMML: Adaptive Multi-Modal Learning for Efficient Video Recognition

Self-supervised Multi-modal Hybrid Fusion Network for Brain Tumor Segmentation

Multi-modal co-attention for drug-target interaction annotation and Its Application to SARS-CoV-2

Code of paper Interact, Embed, and EnlargE (IEEE): Boosting Modality-specific Representations for Multi-Modal Person Re-identification.

Multi-Modal Machine Learning toolkit based on PyTorch.

Multi-Modal Machine Learning toolkit based on PaddlePaddle.

Multi-modal Text Recognition Networks: Interactive Enhancements between Visual and Semantic Features

Owner

Silver

Hunt down social media accounts by username across social networks

Code of U2Fusion: a unified unsupervised image fusion network for multiple image fusion tasks, including multi-modal, multi-exposure and multi-focus image fusion.

Code and pre-trained models for MultiMAE: Multi-modal Multi-task Masked Autoencoders

Tensorflow python implementation of "Learning High Fidelity Depths of Dressed Humans by Watching Social Media Dance Videos"

This was initially the repo for the project of PSYC626@USC of Asaf Mazar, Millad Kassaie and Georgios Chochlakis named "Powered by the Will? Exploring Lay Theories of Behavior Change through Social Media"

Machine learning and Deep learning models, deploy on telegram (the best social media)

The DL Streamer Pipeline Zoo is a catalog of optimized media and media analytics pipelines.

[CVPR'21] Multi-Modal Fusion Transformer for End-to-End Autonomous Driving

Deep RGB-D Saliency Detection with Depth-Sensitive Attention and Automatic Multi-Modal Fusion (CVPR'2021, Oral)

A Multi-modal Model Chinese Spell Checker Released on ACL2021.