L3Cube-MahaCorpus a Marathi monolingual data set scraped from different internet sources.

Last update: Dec 17, 2022

Related tags

Text Data & NLP nlp machine-learning natural-language-processing sentiment-analysis text-classification dataset marathi bert marathi-language marathi-nlp

Overview

L3Cube-MahaCorpus

L3Cube-MahaCorpus a Marathi monolingual data set scraped from different internet sources. We expand the existing Marathi monolingual corpus with 24.8M sentences and 289M tokens. We also present, MahaBERT, MahaAlBERT, and MahaRoBerta all BERT-based masked language models, and MahaFT, the fast text word embeddings both trained on full Marathi corpus with 752M tokens. The evaluation details are mentioned in our paper link

Dataset Statistics

L3Cube-MahaCorpus(full) = L3Cube-MahaCorpus(news) + L3Cube-MahaCorpus(non-news)

Full Marathi Corpus incorporates all existing sources .

Dataset	#tokens(M)	#sentences(M)	Link
L3Cube-MahaCorpus(news)	212	17.6	link
L3Cube-MahaCorpus(non-news)	76.4	7.2	link
L3Cube-MahaCorpus(full)	289	24.8	link
Full Marathi Corpus(all sources)	752	57.2	link

Marathi BERT models and Marathi Fast Text model

The full Marathi Corpus is used to train BERT language models and made available on HuggingFace model hub.

Model	Description	Link
MahaBERT	Base-BERT	link
MahaRoBERTa	RoBERTa	link
MahaAlBERT	AlBERT	link
MahaFT	Fast Text	bin vec

L3CubeMahaSent

L3CubeMahaSent is the largest publicly available Marathi Sentiment Analysis dataset to date. This dataset is made of marathi tweets which are manually labelled. The annotation guidelines are mentioned in our paper link .

Dataset Statistics

This dataset contains a total of 18,378 tweets which are classified into three classes - Positive(1), Negative(-1) and Neutral(0). All tweets are present in their original form, without any preprocessing.

Out of these, 15,864 tweets are considered for splitting them into train(tweets-train.csv), test(tweets-test.csv) and validation(tweets-valid.csv) datasets. This has been done to avoid class imbalance in our dataset.
The remaining 2,514 tweets are also provided in a separate sheet(tweets-extra.csv).

The statistics of the dataset are as follows :

Split	Total tweets	Tweets per class
Train	12114	4038
Test	2250	750
Validation	1500	500

The extra sheet contains 2355 positive and 159 negative tweets. These tweets have not been considered during baseline experiments.

Baseline Experimentations

Two-class(positive,negative) and Three-class(positive,negative,neutral) sentiment analysis / classification was performed on the dataset.

Models

Some of the models used or performing baseline experiments were:

CNN, BiLSTM
- fastText embeddings provided by IndicNLP and Facebook are also used along with the above two models. These embeddings are used in two variations: static and trainable.
BERT based models:
- Multilingual BERT
- IndicBERT

Results

Details of the best performing models are given in the following table:

Model	3-class	2-class
CNN IndicFT trainable	83.24	93.13
BiLSTM IndicFT trainable	82.89	91.80
IndicBERT	84.13	92.93

The fine-tuned IndicBERT model is available on huggingface here . Further details about the dataset and baseline experiments can be found in this paper pdf .

License

L3Cube-MahaCorpus and L3CubeMahaSent is licensed under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License.

Citing

@article{joshi2022l3cube,
  title={L3Cube-MahaCorpus and MahaBERT: Marathi Monolingual Corpus, Marathi BERT Language Models, and Resources},
  author={Joshi, Raviraj},
  journal={arXiv preprint arXiv:2202.01159},
  year={2022}
}

@inproceedings{kulkarni2021l3cubemahasent,
  title={L3CubeMahaSent: A Marathi Tweet-based Sentiment Analysis Dataset},
  author={Kulkarni, Atharva and Mandhane, Meet and Likhitkar, Manali and Kshirsagar, Gayatri and Joshi, Raviraj},
  booktitle={Proceedings of the Eleventh Workshop on Computational Approaches to Subjectivity, Sentiment and Social Media Analysis},
  pages={213--220},
  year={2021}
}

@inproceedings{kulkarni2022experimental,
  title={Experimental evaluation of deep learning models for marathi text classification},
  author={Kulkarni, Atharva and Mandhane, Meet and Likhitkar, Manali and Kshirsagar, Gayatri and Jagdale, Jayashree and Joshi, Raviraj},
  booktitle={Proceedings of the 2nd International Conference on Recent Trends in Machine Learning, IoT, Smart Cities and Applications},
  pages={605--613},
  year={2022},
  organization={Springer}
}

Index different CKAN entities in Solr, not just datasets

ckanext-sitesearch Index different CKAN entities in Solr, not just datasets Requirements This extension requires CKAN 2.9 or higher and Python 3 Featu

3 Dec 2, 2022

A simple Streamlit App to classify swahili news into different categories.

Swahili News Classifier Streamlit App A simple app to classify swahili news into different categories. Installation Install all streamlit requirements

4 May 1, 2022

CredData is a set of files including credentials in open source projects

CredData is a set of files including credentials in open source projects. CredData includes suspicious lines with manual review results and more information such as credential types for each suspicious line. CredData can be used to develop new tools or improve existing tools. Furthermore, using the benchmark result of the CredData, users can choose a proper tool among open source credential scanning tools according to their use case.

19 Sep 7, 2022

This repository contains data used in the NAACL 2021 Paper - Proteno: Text Normalization with Limited Data for Fast Deployment in Text to Speech Systems

Proteno This is the data release associated with the corresponding NAACL 2021 Paper - Proteno: Text Normalization with Limited Data for Fast Deploymen

37 Dec 4, 2022

:mag: End-to-End Framework for building natural language search interfaces to data by utilizing Transformers and the State-of-the-Art of NLP. Supporting DPR, Elasticsearch, HuggingFace’s Modelhub and much more!

Haystack is an end-to-end framework that enables you to build powerful and production-ready pipelines for different search use cases. Whether you want

1.4k Feb 18, 2021

Comments

Data issue in NER training section?
At line 159508 of the train_iob.txt file, sentence 17197.0, there is the following:

या O 17197.0 मंत्राची O 17197.0 देवता O 17197.0 गणपती O 17197.0 ँ O 17197.0 हा O 17197.0 तो O 17197.0 मंत्र O 17197.0

I don't know much about Marathi, to be honest, but to my interpretation this is a Candrabindu mark with no previous character it is marking. I believe that to be an error. Would you confirm, and if so, would you suggest a fix?

Thanks!
opened by AngledLuffa 2

L3Cube-MahaCorpus a Marathi monolingual data set scraped from different internet sources.

Related tags

Overview

L3Cube-MahaCorpus

Dataset Statistics

Marathi BERT models and Marathi Fast Text model

L3CubeMahaSent

Dataset Statistics

Baseline Experimentations

Models

Results

License

Citing

You might also like...

Index different CKAN entities in Solr, not just datasets

A simple Streamlit App to classify swahili news into different categories.

CredData is a set of files including credentials in open source projects

This repository contains data used in the NAACL 2021 Paper - Proteno: Text Normalization with Limited Data for Fast Deployment in Text to Speech Systems

This repository is home to the Optimus data transformation plugins for various data processing needs.

Tools, wrappers, etc... for data science with a concentration on text processing

Data loaders and abstractions for text and NLP

Data loaders and abstractions for text and NLP

:mag: End-to-End Framework for building natural language search interfaces to data by utilizing Transformers and the State-of-the-Art of NLP. Supporting DPR, Elasticsearch, HuggingFace’s Modelhub and much more!

Comments

Data issue in NER training section?

Owner

The Internet Archive Research Assistant - Daily search Internet Archive for new items matching your keywords

justCTF [*] 2020 challenges sources

Jarvis is a simple Chatbot with a GUI capable of chatting and retrieving information and daily news from the internet for it's user.

A collection of Classical Chinese natural language processing models, including Classical Chinese related models and resources on the Internet.

Deduplication is the task to combine different representations of the same real world entity.

Smart discord chatbot integrated with Dialogflow to manage different classrooms and assist in teaching!

Explore different way to mix speech model(wav2vec2, hubert) and nlp model(BART,T5,GPT) together

Continuously update some NLP practice based on different tasks.

Takes a string and puts it through different languages in Google Translate a requested amount of times, returning nonsense.