Using Python to scrape some basic player information from www.premierleague.com and then use Pandas to analyse said data.

Last update: Sep 6, 2021

Related tags

Data Analysis PremiershipPlayerAnalysis

Overview

PremiershipPlayerAnalysis

Using Python to scrape some basic player information from www.premierleague.com and then use Pandas to analyse said data. Note : My understanding is the squad data on this site can change at any time so your results might be different

Improvement : Calculate age to finer degree than just years

The was developed in Jupyter Notebook and this walkthrough willl assume you are doing the same

Once you have ran the scraping

original = pd.DataFrame(playersList) # Convert the data scraped into a Pandas DataFrame 

original.to_csv('premiershipplayers.csv') # Keep a back up of the data to save time later if required 

df2 = original.copy() # Working copy of the DataFrame (just in case) 


df2.info()


   
    
RangeIndex: 578 entries, 0 to 577
Data columns (total 11 columns):
 #   Column       Non-Null Count  Dtype 
---  ------       --------------  ----- 
 0   club         578 non-null    object
 1   name         578 non-null    object
 2   shirtNo      572 non-null    object
 3   nationality  562 non-null    object
 4   dob          562 non-null    object
 5   height       500 non-null    object
 6   weight       474 non-null    object
 7   appearances  578 non-null    object
 8   goals        578 non-null    object
 9   wins         578 non-null    object
 10  losses       578 non-null    object
dtypes: object(11)
memory usage: 49.8+ KB

*** A total of 578 player. ***

6 without shirt number

16 without nationality listed

16 without dob listed

78 without height listed

104 without weight listed

Cleanup Data

Remove spaces and newline from dob, appearances, goals, wins and losses columns
Change type of dob to date

change type of appearances, goals, wins, losses to int

 df2['dob'] = df2['dob'].str.replace('\n','').str.strip(' ')
 df2['appearances'] = df2['appearances'].str.replace('\n','').str.strip(' ')
 df2['goals'] = df2['goals'].str.replace('\n','').str.strip(' ')
 df2['wins'] = df2['wins'].str.replace('\n','').str.strip(' ')
 df2['losses'] = df2['losses'].str.replace('\n','').str.strip(' ')

 # change type of dob, appearances, goals, wins, losses
 from datetime import  date

 df2['dob'] = pd.to_datetime(df2['dob'],format='%d/%m/%Y').dt.date
 df2["appearances"] = pd.to_numeric(df2["appearances"])
 df2["goals"] = pd.to_numeric(df2["goals"])
 df2["wins"] = pd.to_numeric(df2["wins"])
 df2["losses"] = pd.to_numeric(df2["losses"])
 df2['height'] = df2['height'].str[:-2]
 df2["height"] = pd.to_numeric(df2["height"])
 
 
 # Create age column

 today = date.today()

 def age(born):
     if born:
         return today.year - born.year - ((today.month, 
                                       today.day) < (born.month, 
                                                     born.day))
     else:
         return np.nan

 df2['age'] = df2['dob'].apply(age)

10 Oldest Players

    df2.sort_values('age',ascending=False).head(10)

10 Youngest Players

    df2.sort_values('age',ascending=True).head(10)

Squad Sizes

    df2.groupby(['club'])['club'].count().sort_values(ascending=False)

Team's Average Player Age

    plt.ylim([20, 30])
    df2.groupby(['club'])['age'].mean().sort_values(ascending=False).plot.bar()

Burnley appear to not only have one of the highest average player ages but also the owest number of registered players

Top 10 Premiership Appearances

    df2.sort_values('appearances',ascending=False).head(10)

Collective Premiership Appearances per Club

    df2.groupby(['club'])['appearances'].sum().sort_values(ascending=False)

    df2.groupby(['club'])['appearances'].sum().sort_values(ascending=False).plot.bar()

10 Tallest Playes

    df2.sort_values('height',ascending=False).head(10)

10 Shortest Playes

    df2.sort_values('height',ascending=True).head(10)

Nationality totals of Players

    pd.set_option('display.max_rows', 100)
    df.groupby(['nationality'])['club'].count().sort_values(ascending=False)

Nationality totals per club

    pd.set_option('display.max_rows', 500)
    df.groupby(['club','nationality'])['nationality'].count()

You might also like...

Pandas on AWS - Easy integration with Athena, Glue, Redshift, Timestream, QuickSight, Chime, CloudWatchLogs, DynamoDB, EMR, SecretManager, PostgreSQL, MySQL, SQLServer and S3 (Parquet, CSV, JSON and EXCEL).

AWS Data Wrangler Pandas on AWS Easy integration with Athena, Glue, Redshift, Timestream, QuickSight, Chime, CloudWatchLogs, DynamoDB, EMR, SecretMana

3.3k Jan 4, 2023

Using Python to scrape some basic player information from www.premierleague.com and then use Pandas to analyse said data.

Related tags

Overview

PremiershipPlayerAnalysis

Cleanup Data

10 Oldest Players

10 Youngest Players

Squad Sizes

Team's Average Player Age

Burnley appear to not only have one of the highest average player ages but also the owest number of registered players

Top 10 Premiership Appearances

Collective Premiership Appearances per Club

10 Tallest Playes

10 Shortest Playes

Nationality totals of Players

Nationality totals per club

You might also like...

Pandas on AWS - Easy integration with Athena, Glue, Redshift, Timestream, QuickSight, Chime, CloudWatchLogs, DynamoDB, EMR, SecretManager, PostgreSQL, MySQL, SQLServer and S3 (Parquet, CSV, JSON and EXCEL).

Statistical package in Python based on Pandas

Projeto para realizar o RPA Challenge . Utilizando Python e as bibliotecas Selenium e Pandas.

Python utility to extract differences between two pandas dataframes.

Pandas-based utility to calculate weighted means, medians, distributions, standard deviations, and more.

Pandas and Dask test helper methods with beautiful error messages.

Pandas and Spark DataFrame comparison for humans

An extension to pandas dataframes describe function.

Create HTML profiling reports from pandas DataFrame objects

Owner

A set of tools to analyse the output from TraDIS analyses

A Pythonic introduction to methods for scaling your data science and machine learning work to larger datasets and larger models, using the tools and APIs you know and love from the PyData stack (such as numpy, pandas, and scikit-learn).

A data analysis using python and pandas to showcase trends in school performance.

Created covid data pipeline using PySpark and MySQL that collected data stream from API and do some processing and store it into MYSQL database.

Reading streams of Twitter data, save them to Kafka, then process with Kafka Stream API and Spark Streaming

Hatchet is a Python-based library that allows Pandas dataframes to be indexed by structured tree and graph data.

NumPy and Pandas interface to Big Data

Finds, downloads, parses, and standardizes public bikeshare data into a standard pandas dataframe format

A powerful data analysis package based on mathematical step functions. Strongly aligned with pandas.

Calculate multilateral price indices in Python (with Pandas and PySpark).