Implementing SimCSE(paper, official repository) using TensorFlow 2 and KR-BERT.

Jeong Ukjae

Last update: Dec 12, 2022

Related tags

Text Data & NLP paper-implementations

Overview

KR-BERT-SimCSE

Implementing SimCSE(paper, official repository) using TensorFlow 2 and KR-BERT.

Training

Unsupervised

python train_unsupervised.py --mixed_precision

I used Korean Wikipedia Corpus that is divided into sentences in advance. (Check out tfds-korean catalog page for details)

Settings
- KR-BERT character
- peak learning rate 3e-5
- batch size 64
- Total steps: 25,000
- 0.05 warmup rate, and linear decay learning rate scheduler
- temperature 0.05
- evalaute on KLUE STS and KorSTS every 250 steps
- max sequence length 64
- Use pooled outputs for training, and [CLS] token's representations for inference

The hyperparameters were not tuned and mostly followed the values in the paper.

Supervised

python train_supervised.py --mixed_precision

I used KorNLI for supervised training. (Check out tfds-korean catalog page)

Settings
- KR-BERT character
- batch size 128
- epoch 3
- peak learning rate 5e-5
- 0.05 warmup rate, and linear decay learning rate scheduler
- temperature 0.05
- evalaute on KLUE STS and KorSTS every 125 steps
- max sequence length 48
- Use pooled outputs for training, and [CLS] token's representations for inference

The hyperparameters were not tuned and mostly followed the values in the paper.

Results

KorSTS (dev set results)

model			100 X Spearman correlation
KR-BERT base SimCSE	unsupervised	bi encoding	79.99
KR-BERT base SimCSE-supervised	trained on KorNLI	bi encoding	84.88

SRoBERTa base*	unsupervised	bi encoding	63.34
SRoBERTa base*	trained on KorNLI	bi encoding	76.48
SRoBERTa base*	trained on KorSTS	bi encoding	83.68
SRoBERTa base*	trained on KorNLI -> KorSTS	bi encoding	83.54

SRoBERTa large*	trained on KorNLI	bi encoding	77.95
SRoBERTa large*	trained on KorSTS	bi encoding	84.74
SRoBERTa large*	trained on KorNLI -> KorSTS	bi encoding	84.21

*: results from Ham et al., 2020.

KorSTS (test set results)

model			100 X Spearman correlation
KR-BERT base SimCSE	unsupervised	bi encoding	73.25
KR-BERT base SimCSE-supervised	trained on KorNLI	bi encoding	80.72

SRoBERTa base*	unsupervised	bi encoding	48.96
SRoBERTa base*	trained on KorNLI	bi encoding	74.19
SRoBERTa base*	trained on KorSTS	bi encoding	78.94
SRoBERTa base*	trained on KorNLI -> KorSTS	bi encoding	80.29

SRoBERTa large*	trained on KorNLI	bi encoding	75.46
SRoBERTa large*	trained on KorSTS	bi encoding	79.55
SRoBERTa large*	trained on KorNLI -> KorSTS	bi encoding	80.49

SRoBERTa base*	trained on KorSTS	cross encoding	83.00
SRoBERTa large*	trained on KorSTS	cross encoding	85.27

*: results from Ham et al., 2020.

KLUE STS (dev set results)

model			100 X Pearson's correlation
KR-BERT base SimCSE	unsupervised	bi encoding	74.45
KR-BERT base SimCSE-supervised	trained on KorNLI	bi encoding	79.42

KR-BERT base*	supervised	cross encoding	87.50

*: results from Park et al., 2021.

References

@misc{gao2021simcse,
    title={SimCSE: Simple Contrastive Learning of Sentence Embeddings},
    author={Tianyu Gao and Xingcheng Yao and Danqi Chen},
    year={2021},
    eprint={2104.08821},
    archivePrefix={arXiv},
    primaryClass={cs.CL}
}

@misc{ham2020kornli,
    title={KorNLI and KorSTS: New Benchmark Datasets for Korean Natural Language Understanding},
    author={Jiyeon Ham and Yo Joong Choe and Kyubyong Park and Ilji Choi and Hyungjoon Soh},
    year={2020},
    eprint={2004.03289},
    archivePrefix={arXiv},
    primaryClass={cs.CL}
}

@misc{park2021klue,
    title={KLUE: Korean Language Understanding Evaluation},
    author={Sungjoon Park and Jihyung Moon and Sungdong Kim and Won Ik Cho and Jiyoon Han and Jangwon Park and Chisung Song and Junseong Kim and Yongsook Song and Taehwan Oh and Joohong Lee and Juhyun Oh and Sungwon Lyu and Younghoon Jeong and Inkwon Lee and Sangwoo Seo and Dongjun Lee and Hyunwoo Kim and Myeonghwa Lee and Seongbo Jang and Seungwon Do and Sunkyoung Kim and Kyungtae Lim and Jongwon Lee and Kyumin Park and Jamin Shin and Seonghyun Kim and Lucy Park and Alice Oh and Jung-Woo Ha and Kyunghyun Cho},
    year={2021},
    eprint={2105.09680},
    archivePrefix={arXiv},
    primaryClass={cs.CL}
}

You might also like...

Natural language processing summarizer using 3 state of the art Transformer models: BERT, GPT2, and T5

NLP-Summarizer Natural language processing summarizer using 3 state of the art Transformer models: BERT, GPT2, and T5 This project aimed to provide in

1 Feb 7, 2022

This repository contains the official release of the model "BanglaBERT" and associated downstream finetuning code and datasets introduced in the paper titled "BanglaBERT: Combating Embedding Barrier in Multilingual Models for Low-Resource Language Understanding".

BanglaBERT This repository contains the official release of the model "BanglaBERT" and associated downstream finetuning code and datasets introduced i

197 Dec 25, 2022

TunBERT is the first release of a pre-trained BERT model for the Tunisian dialect using a Tunisian Common-Crawl-based dataset.

TunBERT is the first release of a pre-trained BERT model for the Tunisian dialect using a Tunisian Common-Crawl-based dataset. TunBERT was applied to three NLP downstream tasks: Sentiment Analysis (SA), Tunisian Dialect Identification (TDI) and Reading Comprehension Question-Answering (RCQA)

72 Dec 9, 2022

Using Bert as the backbone model for lime, designed for NLP task explanation (sentence pair text classification task)

Lime Comparing deep contextualized model for sentences highlighting task. In addition, take the classic explanation model "LIME" with bert-base model

2 Jan 18, 2022

Multilingual Emotion classification using BERT (fine-tuning). Published at the WASSA workshop (ACL2022).

XLM-EMO: Multilingual Emotion Prediction in Social Media Text Abstract Detecting emotion in text allows social and computational scientists to study h

35 Sep 17, 2022

Kashgari is a production-level NLP Transfer learning framework built on top of tf.keras for text-labeling and text-classification, includes Word2Vec, BERT, and GPT2 Language Embedding.

Kashgari Overview | Performance | Installation | Documentation | Contributing 🎉 🎉 🎉 We released the 2.0.0 version with TF2 Support. 🎉 🎉 🎉 If you

2.3k Dec 29, 2022

Kashgari is a production-level NLP Transfer learning framework built on top of tf.keras for text-labeling and text-classification, includes Word2Vec, BERT, and GPT2 Language Embedding.

Kashgari Overview | Performance | Installation | Documentation | Contributing 🎉 🎉 🎉 We released the 2.0.0 version with TF2 Support. 🎉 🎉 🎉 If you

2k Feb 9, 2021

PhoNLP: A BERT-based multi-task learning toolkit for part-of-speech tagging, named entity recognition and dependency parsing

PhoNLP is a multi-task learning model for joint part-of-speech (POS) tagging, named entity recognition (NER) and dependency parsing. Experiments on Vietnamese benchmark datasets show that PhoNLP produces state-of-the-art results, outperforming a single-task learning approach that fine-tunes the pre-trained Vietnamese language model PhoBERT for each task independently.

109 Dec 2, 2022

🛸 Use pretrained transformers like BERT, XLNet and GPT-2 in spaCy

spacy-transformers: Use pretrained transformers like BERT, XLNet and GPT-2 in spaCy This package provides spaCy components and architectures to use tr

1.2k Jan 8, 2023

Implementing SimCSE(paper, official repository) using TensorFlow 2 and KR-BERT.

Related tags

Overview

KR-BERT-SimCSE

Training

Unsupervised

Supervised

Results

KorSTS (dev set results)

KorSTS (test set results)

KLUE STS (dev set results)

References

You might also like...

Natural language processing summarizer using 3 state of the art Transformer models: BERT, GPT2, and T5

This repository contains the official release of the model "BanglaBERT" and associated downstream finetuning code and datasets introduced in the paper titled "BanglaBERT: Combating Embedding Barrier in Multilingual Models for Low-Resource Language Understanding".

TunBERT is the first release of a pre-trained BERT model for the Tunisian dialect using a Tunisian Common-Crawl-based dataset.

Using Bert as the backbone model for lime, designed for NLP task explanation (sentence pair text classification task)

Multilingual Emotion classification using BERT (fine-tuning). Published at the WASSA workshop (ACL2022).

Kashgari is a production-level NLP Transfer learning framework built on top of tf.keras for text-labeling and text-classification, includes Word2Vec, BERT, and GPT2 Language Embedding.

Kashgari is a production-level NLP Transfer learning framework built on top of tf.keras for text-labeling and text-classification, includes Word2Vec, BERT, and GPT2 Language Embedding.

PhoNLP: A BERT-based multi-task learning toolkit for part-of-speech tagging, named entity recognition and dependency parsing

🛸 Use pretrained transformers like BERT, XLNet and GPT-2 in spaCy

Owner

Jeong Ukjae

SimCSE: Simple Contrastive Learning of Sentence Embeddings

MILES is a multilingual text simplifier inspired by LSBert - A BERT-based lexical simplification approach proposed in 2018. Unlike LSBert, MILES uses the bert-base-multilingual-uncased model, as well as simple language-agnostic approaches to complex word identification (CWI) and candidate ranking.

VD-BERT: A Unified Vision and Dialog Transformer with BERT

LV-BERT: Exploiting Layer Variety for BERT (Findings of ACL 2021)

Pytorch-version BERT-flow: One can apply BERT-flow to any PLM within Pytorch framework.

Python module (C extension and plain python) implementing Aho-Corasick algorithm

Python module (C extension and plain python) implementing Aho-Corasick algorithm

A Python package implementing a new model for text classification with visualization tools for Explainable AI :octocat:

A framework for implementing federated learning

Code of paper: A Recurrent Vision-and-Language BERT for Navigation