Vaex library for Big Data Analytics of an Airline dataset

Nikolas Petrou

Last update: Feb 13, 2022

Related tags

Data Analysis visualization python data-science machine-learning big-data exploratory-data-analysis jupyter-notebook regression eda big-data-analytics vaex us-airline-dataset

Overview

Vaex-Big-Data-Analytics-for-Airline-data

A Python notebook (ipynb) created in Jupyter Notebook, which utilizes the Vaex library for Big Data Analytics of an Airline dataset.

Author: Nikolas Petrou, MSc in Data Science

Overview

The main part of the work focuses on the exploration a big dataset of 17 GB. Specifically, the dataset contains information on flights within the United States between 1988 and 2018. It can be directly downloaded from: vaex.s3.us-east-2.amazonaws.com.

In addition, in this project the Out-of-Core DataFrames Python library Vaex is employed, in order to visualize, explore acalculate statistics of this big tabular dataset.

The goal of this project is to utilize Vaex to perform an Exploratory Data Analysis (EDA), as well as to predict the arrival delay of a flight using Machine Learning models (regression task).

What is Vaex and why Vaex?

Vaex is a Python library for lazy Out-of-Core DataFrames (similar to Pandas), to visualize and explore big tabular datasets. It can calculate statistics such as mean, sum, count, standard deviation etc, on an N-dimensional grid up to a billion (10^9) objects/rows per second. Visualization is done using histograms, density plots and 3d volume rendering, allowing interactive exploration of big data. Furthermore, Vaex provides wrappers to powerful libraries for predictive models (e.g. Scikit-learn, xgboost) and make them work efficiently with Vaex. Vaex does implement a variety of standard data transformers (e.g. PCA, numerical scalers, categorical encoders) and a very efficient KMeans algorithm that take full advantage. Finally, Vaex uses memory mapping, a zero memory copy policy, and lazy computations for best performance (no memory wasted).

Advantage of using Vaex over using Pandas with a more powerful machine

Switching to a more powerful machine (with more RAM and/or better CPU) may solve some memory issues, but still, Pandas will only use one out of the 32 cores of your fancy machine. With Vaex, all operations are out of the core and executed in parallel and lazily evaluated, allowing for crunching through a billion-row dataset effortlessly.

Data

The dataset has a relatively big size (17 GB), and contains information on flights within the United States between 1988 and 2018. It can be directly downloaded from: vaex.s3.us-east-2.amazonaws.com

Each row-record of the dataset represents an individual flight. Specifically, each record contains information of the airline (UniqueCarrier), airports (origin airport, destination airport) and flight level information such as time schedule (day of week, day of month, month, year), flight distance, departure time and delay, arrival time and delay.

You might also like...

Elementary is an open-source data reliability framework for modern data teams. The first module of the framework is data lineage.

Data lineage made simple, reliable, and automated. Effortlessly track the flow of data, understand dependencies and analyze impact. Features Visualiza

898 Jan 9, 2023

Retail-Sim is python package to easily create synthetic dataset of retaile store.

Retailer's Sale Data Simulation Retail-Sim is python package to easily create synthetic dataset of retaile store. Simulation Model Simulator consists

7 Sep 30, 2022

A python package which can be pip installed to perform statistics and visualize binomial and gaussian distributions of the dataset

GBiStat package A python package to assist programmers with data analysis. This package could be used to plot : Binomial Distribution of the dataset p

4 Oct 17, 2022

Pipeline and Dataset helpers for complex algorithm evaluation.

tpcp - Tiny Pipelines for Complex Problems A generic way to build object-oriented datasets and algorithm pipelines and tools to evaluate them pip inst

Machine Learning and Data Analytics Lab FAU

3 Dec 7, 2022

This is a python script to navigate and extract the FSD50K dataset

FSD50K navigator This is a script I use to navigate the sound dataset from FSK50K.

2 Nov 23, 2021

This is an example of how to automate Ridit Analysis for a dataset with large amount of questions and many item attributes

1 Nov 17, 2021

Vaex library for Big Data Analytics of an Airline dataset

Related tags

Overview

Vaex-Big-Data-Analytics-for-Airline-data

Overview

What is Vaex and why Vaex?

Advantage of using Vaex over using Pandas with a more powerful machine

Data

You might also like...

Elementary is an open-source data reliability framework for modern data teams. The first module of the framework is data lineage.

Retail-Sim is python package to easily create synthetic dataset of retaile store.

A python package which can be pip installed to perform statistics and visualize binomial and gaussian distributions of the dataset

Pipeline and Dataset helpers for complex algorithm evaluation.

This is a python script to navigate and extract the FSD50K dataset

This is an example of how to automate Ridit Analysis for a dataset with large amount of questions and many item attributes

Python dataset creator to construct datasets composed of OpenFace extracted features and Shimmer3 GSR+ Sensor datas

For making Tagtog annotation into csv dataset

pyETT: Python library for Eleven VR Table Tennis data

Owner

Nikolas Petrou

Tuplex is a parallel big data processing framework that runs data science pipelines written in Python at the speed of compiled code

A Big Data ETL project in PySpark on the historical NYC Taxi Rides data

NumPy and Pandas interface to Big Data

The official repository for ROOT: analyzing, storing and visualizing big data, scientifically

BigDL - Evaluate the performance of BigDL (Distributed Deep Learning on Apache Spark) in big data analysis problems

ForecastGA is a Python tool to forecast Google Analytics data using several popular time series models.

Tokyo 2020 Paralympics, Analytics

Mortgage-loan-prediction - Show how to perform advanced Analytics and Machine Learning in Python using a full complement of PyData utilities

Exploratory Data Analysis for Employee Retention Dataset

Amundsen is a metadata driven application for improving the productivity of data analysts, data scientists and engineers when interacting with data.