Skip to content

Latest commit

 

History

5 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 

Repository files navigation

Data Science Job Salaries — Analysis and Prediction

Overview

This project analyzes global data science job salaries to identify factors influencing compensation across different regions, experience levels, company sizes, and remote work arrangements. Through exploratory data analysis (EDA), feature engineering, and machine learning modeling, the study provides insights into salary trends and key drivers in the data science job market.

Dataset Description

The dataset contains comprehensive information on job roles in the data science field, covering multiple years, locations, and employment types.

Column Description
work_year The year the salary was paid.
experience_level The experience level in the job during the year:
• EN – Entry-level / Junior
• MI – Mid-level / Intermediate
• SE – Senior-level / Expert
• EX – Executive-level / Director
employment_type The type of employment:
• PT – Part-time
• FT – Full-time
• CT – Contract
• FL – Freelance
job_title The role worked in during the year.
salary The total gross salary amount paid.
salary_currency The currency of the salary paid (ISO 4217 format).
salary_in_usd The salary converted to USD for standard comparison.
employee_residence Employee's country of residence (ISO 3166 format).
remote_ratio Proportion of remote work:
• 0 – No remote work
• 50 – Partially remote
• 100 – Fully remote
company_location The country of the employer’s main office (ISO 3166 format).
company_size Average company size:
• S – <50 employees (small)
• M – 50–250 employees (medium)
• L – >250 employees (large)

Data Preparation and Feature Engineering

To improve model interpretability and accuracy, several transformations were applied:

  • Creation of a binary flag for remote work, distinguishing fully remote positions from others.
  • Log transformation of salaries to reduce skewness in the salary distribution.
  • Comparison feature (same_country) to identify whether the employee and company are based in the same country.
  • Ordinal encoding for experience level and employment type to reflect hierarchy and relationships.
  • Label encoding for categorical variables such as job title, company size, and locations.

Modeling and Performance

Multiple machine learning algorithms were tested, including CatBoost, XGBoost, LightGBM, and Random Forest, with hyperparameter optimization via Optuna. A Weighted Ensemble Model delivered the best results.

Rank Model MAE RMSE R² Score
1 Weighted Ensemble 24,292.65 31,495.21 0.6276
2 CatBoost (Optuna Tuned) 24,247.52 31,883.74 0.6184
3 CatBoost Top 24,759.72 32,334.35 0.6075
4 CatBoost (Deeper Search) 24,686.19 32,810.12 0.5959
5 CatBoost (Tuned) 25,209.29 32,834.65 0.5952
6 Stacking Ensemble 25,669.85 33,169.46 0.5869
7 LightGBM 25,862.07 33,280.96 0.5842
8 XGBoost (Tuned) 25,772.36 33,375.02 0.5818
9 Random Forest (Tuned) 26,478.87 34,060.20 0.5645

SHAP Analysis — Explainable AI

To understand how features influenced the model’s predictions, SHAP (SHapley Additive exPlanations) was applied.

Key Insights

  • Location factors - employee residence and company location were the most influential features.
  • Experience level - higher experience levels consistently correlated with higher salaries.
  • Job title - specialized roles commanded premium pay.
  • Company size - larger organizations generally offered better compensation.
  • Remote work - fully remote positions showed slightly higher average salaries.
  • Work year - salaries have shown a steady upward trend over time.

In summary, location, experience level, and job title are the primary salary drivers, with secondary influences from company size, work year, and remote work opportunities.

Technologies Used

  • Python
  • Pandas, NumPy, Matplotlib, Seaborn
  • Scikit-learn, Optuna, SHAP
  • CatBoost, XGBoost, LightGBM
  • Jupyter Notebook