PLStream: A Framework for Fast Polarity Labelling of Massive Data Streams

Last update: Aug 2, 2022

Related tags

Data Analysis PLStream

Overview

PLStream: A Framework for Fast Polarity Labelling of Massive Data Streams

Motivation

When dataset freshness is critical, the annotating of high speed unlabelled data streams becomes critical but remains an open problem.
We propose PLStream, a novel Apache Flink-based framework for fast polarity labelling of massive data streams, like Twitter tweets or online product reviews.

Environment Requirements

relative python packages are summerized in requirements.txt

Flink v1.13
Python 3.7
Java 8

DataSource

Dataset quick access on https://course.fast.ai/datasets#nlp

Tweets

1.6 million labeled Tweets:
Source:Sentiment140

Yelp Reviews

280,000 training and 19,000 test samples in each polarity
Source:Yelp Review Polarity

Amazon Reviews

1,800,000 training and 200,000 testing samples in each polarity
Source:Amazon product review polarity

Quick Start

quick try PLStream on yelp review dataset

Data Prepare

cd PLStream
weget https://s3.amazonaws.com/fast-ai-nlp/yelp_review_polarity_csv.tgz
tar zxvf yelp_review_polarity_csv.tgz
mv yelp_review_polarity_csv/train.csv train.csv

1. Install required environment of PLStream

please make sure Environment Requirements mentioned above is ready.

pip install -r requirements.txt

2. Start Redis-Server in a terminal

redis-server

3. Run PLStream

python PLStream.py

The outputs' form is "original text" + "label" + "@@@@":
With help of a split("@@@@") function we can further reorganize the labelled dataset.

Optional

to see the labelling accuracy, simply run: python PLStream_acc.py

You might also like...

A data parser for the internal syncing data format used by Fog of World.

A data parser for the internal syncing data format used by Fog of World. The parser is not designed to be a well-coded library with good performance, it is more like a demo for showing the data structure.

40 Dec 12, 2022

Functional Data Analysis, or FDA, is the field of Statistics that analyses data that depend on a continuous parameter.

Functional Data Analysis Python package

Grupo de Aprendizaje Automático - Universidad Autónoma de Madrid

184 Dec 27, 2022

Fancy data functions that will make your life as a data scientist easier.

WhiteBox Utilities Toolkit: Tools to make your life easier Fancy data functions that will make your life as a data scientist easier. Installing To ins

3 Oct 3, 2022

A Big Data ETL project in PySpark on the historical NYC Taxi Rides data

Processing NYC Taxi Data using PySpark ETL pipeline Description This is an project to extract, transform, and load large amount of data from NYC Taxi

2 Dec 12, 2021

Created covid data pipeline using PySpark and MySQL that collected data stream from API and do some processing and store it into MYSQL database.

2 Nov 20, 2021

Utilize data analytics skills to solve real-world business problems using Humana’s big data

Humana-Mays-2021-HealthCare-Analytics-Case-Competition- The goal of the project is to utilize data analytics skills to solve real-world business probl

1 Dec 27, 2021

Python data processing, analysis, visualization, and data operations

Python This is a Python data processing, analysis, visualization and data operations of the source code warehouse, book ISBN: 9787115527592 Descriptio

1 Jan 16, 2022

PrimaryBid - Transform application Lifecycle Data and Design and ETL pipeline architecture for ingesting data from multiple sources to redshift

Transform application Lifecycle Data and Design and ETL pipeline architecture for ingesting data from multiple sources to redshift This project is composed of two parts: Part1 and Part2

1 Jan 19, 2022

Demonstrate the breadth and depth of your data science skills by earning all of the Databricks Data Scientist credentials

Data Scientist Learning Plan Demonstrate the breadth and depth of your data science skills by earning all of the Databricks Data Scientist credentials

27 Nov 1, 2022

PLStream: A Framework for Fast Polarity Labelling of Massive Data Streams

Related tags

Overview

PLStream: A Framework for Fast Polarity Labelling of Massive Data Streams

Motivation

Environment Requirements

DataSource

Tweets

Yelp Reviews

Amazon Reviews

Quick Start

Data Prepare

1. Install required environment of PLStream

2. Start Redis-Server in a terminal

3. Run PLStream

Optional

You might also like...

A data parser for the internal syncing data format used by Fog of World.

Functional Data Analysis, or FDA, is the field of Statistics that analyses data that depend on a continuous parameter.

Fancy data functions that will make your life as a data scientist easier.

A Big Data ETL project in PySpark on the historical NYC Taxi Rides data

Created covid data pipeline using PySpark and MySQL that collected data stream from API and do some processing and store it into MYSQL database.

Utilize data analytics skills to solve real-world business problems using Humana’s big data

Python data processing, analysis, visualization, and data operations

PrimaryBid - Transform application Lifecycle Data and Design and ETL pipeline architecture for ingesting data from multiple sources to redshift

Demonstrate the breadth and depth of your data science skills by earning all of the Databricks Data Scientist credentials

Owner

Reading streams of Twitter data, save them to Kafka, then process with Kafka Stream API and Spark Streaming

Elementary is an open-source data reliability framework for modern data teams. The first module of the framework is data lineage.

Tuplex is a parallel big data processing framework that runs data science pipelines written in Python at the speed of compiled code

A collection of robust and fast processing tools for parsing and analyzing web archive data.

Python package to transfer data in a fast, reliable, and packetized form.

Amundsen is a metadata driven application for improving the productivity of data analysts, data scientists and engineers when interacting with data.

Fast, flexible and easy to use probabilistic modelling in Python.

Sensitivity Analysis Library in Python (Numpy). Contains Sobol, Morris, Fractional Factorial and FAST methods.

🧪 Panel-Chemistry - exploratory data analysis and build powerful data and viz tools within the domain of Chemistry using Python and HoloViz Panel.

fds is a tool for Data Scientists made by DAGsHub to version control data and code at once.