Exploratory Data Analysis for Employee Retention Dataset

kana sudheer reddy

Last update: Oct 1, 2021

Related tags

Data Analysis employee-managment-system

Overview

Exploratory Data Analysis for Employee Retention Dataset

Employee turn-over is a very costly problem for companies.
The cost of replacing an employee if often larger than 100K USD, taking into account the time spent to interview and find a replacement, placement fees, sign-on bonuses and the loss of productivity for several months.
It is only natural then that data science has started being applied to this area.
Understanding why and when employees are most likely to leave can lead to actions to improve employee retention as well as planning new hiring in advance. This application of DS is sometimes called people analytics or people data science
We got employee data from a few companies. We have data about all employees who joined from 2011/01/24 to 2015/12/13. For each employee, we also know if they are still at the company as of 2015/12/13 or they have quit.
Beside that, we have general info about the employee, such as avg salary during her tenure, dept, and yrs of experience.

Goal:

In this challenge, you have a data set with info about the employees and have to predict when employees are going to quit by understanding the main drivers of employee churn.

Assume, for each company, that the headcount starts from zero on 2011/01/23. Estimate employee headcount, for each company, on each day, from 2011/01/24 to 2015/12/13. That is, if by 2012/03/02 2000 people have joined company 1 and 1000 of them have already quit, then company headcount on 2012/03/02 for company 1 would be 1000.
You should create a table with 3 columns: day, employee_headcount, company_id. What are the main factors that drive employee churn? Do they make sense? Explain your findings.
If you could add to this data set just one variable that could help explain employee churn, what would that be?

Data: (data/employee_retention_data.csv)

Columns:

employee_id : id of the employee. Unique by employee per company
company_id : company id.
dept : employee dept
seniority : number of yrs of work experience when hired
salary: avg yearly salary of the employee during her tenure within the company
join_date: when the employee joined the company, it can only be between 2011/01/24 and 2015/12/13
quit_date: when the employee left her job (if she is still employed as of 2015/12/13, this field is NA)

Question 1

Function that returns a list of the names of categorical variables

Define a function with name get_categorical_variables
Pass dataframe as parameter (Read csv file and convert it into pandas dataframe)
Return list of all categorical fields available.

Question 2

Function that returns the list of the names of numeric variables

Define a function with name get_numerical_variables
Pass dataframe as parameter (Read csv file and convert it into pandas dataframe)
Return list of all numerical fields available.

Question 3

Function that returns, for numeric variables, mean, median, 25, 50, 75th percentile

Define a function with name get_numerical_variables_percentile
Pass dataframe as parameter (Read csv file and convert it into pandas dataframe)
Return dataframe with following columns:
- variable name
- mean
- median
- 25th percentile
- 50th percentile
- 75th percentile

Question 4

For categorical variables, get modes

Define a function with name get_categorical_variables_modes
Pass dataframe as parameter (Read csv file and convert it into pandas dataframe)
Return dict object with following keys:
- converted
- country
- new_user
- source

Question 5

For each column, list the count of missing values

Define a function with name get_missing_values_count
Pass dataframe as parameter (Read csv file and convert it into pandas dataframe)
Return dataframe with following columns:
- var_name
- missing_value_count

Question 6

Plot histograms using different subplots of all the numerical values in a single plot

Define a function with name plot_histogram_with_numerical_values
Pass dataframe and list of columns you want to plot as parameter
Plot the graph
Add column names as plot names (In case you dont understand this please connect with instructor)
Change the histogram colour to yellow
Fit a normal curve on those histograms (In case you dont understand this please connect with instructor)

The OHSDI OMOP Common Data Model allows for the systematic analysis of healthcare observational databases.

14 Jan 2, 2023

Repositori untuk menyimpan material Long Course STMKGxHMGI tentang Geophysical Python for Seismic Data Analysis

Long Course "Geophysical Python for Seismic Data Analysis" Instruktur: Dr.rer.nat. Wiwit Suryanto, M.Si Dipersiapkan oleh: Anang Sahroni Waktu: Sesi 1

0 Dec 4, 2021

A data analysis using python and pandas to showcase trends in school performance.

A data analysis using python and pandas to showcase trends in school performance. A data analysis to showcase trends in school performance using Panda

0 Sep 7, 2021

Toolchest provides APIs for scientific and bioinformatic data analysis.

Toolchest Python Client Toolchest provides APIs for scientific and bioinformatic data analysis. It allows you to abstract away the costliness of runni

11 Jun 30, 2022

Additional tools for particle accelerator data analysis and machine information

PyLHC Tools This package is a collection of useful scripts and tools for the Optics Measurements and Corrections group (OMC) at CERN. Documentation Au

3 Apr 13, 2022

A powerful data analysis package based on mathematical step functions. Strongly aligned with pandas.

The leading use-case for the staircase package is for the creation and analysis of step functions. Pretty exciting huh. But don't hit the close button

48 Dec 21, 2022

A collection of learning outcomes data analysis using Python and SQL, from DQLab.

Data Analyst with PYTHON Data Analyst berperan dalam menghasilkan analisa data serta mempresentasikan insight untuk membantu proses pengambilan keputu

6 Oct 11, 2022

DaDRA (day-druh) is a Python library for Data-Driven Reachability Analysis.

DaDRA (day-druh) is a Python library for Data-Driven Reachability Analysis. The main goal of the package is to accelerate the process of computing estimates of forward reachable sets for nonlinear dynamical systems.

2 Nov 8, 2021

Python-based Space Physics Environment Data Analysis Software

pySPEDAS pySPEDAS is an implementation of the SPEDAS framework for Python. The Space Physics Environment Data Analysis Software (SPEDAS) framework is

98 Dec 22, 2022

Exploratory Data Analysis for Employee Retention Dataset

Related tags

Overview

Exploratory Data Analysis for Employee Retention Dataset

You might also like...

The OHSDI OMOP Common Data Model allows for the systematic analysis of healthcare observational databases.

Repositori untuk menyimpan material Long Course STMKGxHMGI tentang Geophysical Python for Seismic Data Analysis

A data analysis using python and pandas to showcase trends in school performance.

Toolchest provides APIs for scientific and bioinformatic data analysis.

Additional tools for particle accelerator data analysis and machine information

A powerful data analysis package based on mathematical step functions. Strongly aligned with pandas.

A collection of learning outcomes data analysis using Python and SQL, from DQLab.

DaDRA (day-druh) is a Python library for Data-Driven Reachability Analysis.

Python-based Space Physics Environment Data Analysis Software

Owner

kana sudheer reddy

Flenser is a simple, minimal, automated exploratory data analysis tool.

Exploratory data analysis

Statistical Analysis 📈 focused on statistical analysis and exploration used on various data sets for personal and professional projects.

This is an example of how to automate Ridit Analysis for a dataset with large amount of questions and many item attributes

A set of functions and analysis classes for solvation structure analysis

Functional Data Analysis, or FDA, is the field of Statistics that analyses data that depend on a continuous parameter.

Python data processing, analysis, visualization, and data operations

Visions provides an extensible suite of tools to support common data analysis operations

Tablexplore is an application for data analysis and plotting built in Python using the PySide2/Qt toolkit.

Universal data analysis tools for atmospheric sciences