This project analyzes global data science job salaries to identify factors influencing compensation across different regions, experience levels, company sizes, and remote work arrangements. Through exploratory data analysis (EDA), feature engineering, and machine learning modeling, the study provides insights into salary trends and key drivers in the data science job market.
The dataset contains comprehensive information on job roles in the data science field, covering multiple years, locations, and employment types.
| Column | Description |
|---|---|
| work_year | The year the salary was paid. |
| experience_level | The experience level in the job during the year: • EN – Entry-level / Junior • MI – Mid-level / Intermediate • SE – Senior-level / Expert • EX – Executive-level / Director |
| employment_type | The type of employment: • PT – Part-time • FT – Full-time • CT – Contract • FL – Freelance |
| job_title | The role worked in during the year. |
| salary | The total gross salary amount paid. |
| salary_currency | The currency of the salary paid (ISO 4217 format). |
| salary_in_usd | The salary converted to USD for standard comparison. |
| employee_residence | Employee's country of residence (ISO 3166 format). |
| remote_ratio | Proportion of remote work: • 0 – No remote work • 50 – Partially remote • 100 – Fully remote |
| company_location | The country of the employer’s main office (ISO 3166 format). |
| company_size | Average company size: • S – <50 employees (small) • M – 50–250 employees (medium) • L – >250 employees (large) |
To improve model interpretability and accuracy, several transformations were applied:
- Creation of a binary flag for remote work, distinguishing fully remote positions from others.
- Log transformation of salaries to reduce skewness in the salary distribution.
- Comparison feature (
same_country) to identify whether the employee and company are based in the same country. - Ordinal encoding for experience level and employment type to reflect hierarchy and relationships.
- Label encoding for categorical variables such as job title, company size, and locations.
Multiple machine learning algorithms were tested, including CatBoost, XGBoost, LightGBM, and Random Forest, with hyperparameter optimization via Optuna. A Weighted Ensemble Model delivered the best results.
| Rank | Model | MAE | RMSE | R² Score |
|---|---|---|---|---|
| 1 | Weighted Ensemble | 24,292.65 | 31,495.21 | 0.6276 |
| 2 | CatBoost (Optuna Tuned) | 24,247.52 | 31,883.74 | 0.6184 |
| 3 | CatBoost Top | 24,759.72 | 32,334.35 | 0.6075 |
| 4 | CatBoost (Deeper Search) | 24,686.19 | 32,810.12 | 0.5959 |
| 5 | CatBoost (Tuned) | 25,209.29 | 32,834.65 | 0.5952 |
| 6 | Stacking Ensemble | 25,669.85 | 33,169.46 | 0.5869 |
| 7 | LightGBM | 25,862.07 | 33,280.96 | 0.5842 |
| 8 | XGBoost (Tuned) | 25,772.36 | 33,375.02 | 0.5818 |
| 9 | Random Forest (Tuned) | 26,478.87 | 34,060.20 | 0.5645 |
To understand how features influenced the model’s predictions, SHAP (SHapley Additive exPlanations) was applied.
- Location factors - employee residence and company location were the most influential features.
- Experience level - higher experience levels consistently correlated with higher salaries.
- Job title - specialized roles commanded premium pay.
- Company size - larger organizations generally offered better compensation.
- Remote work - fully remote positions showed slightly higher average salaries.
- Work year - salaries have shown a steady upward trend over time.
In summary, location, experience level, and job title are the primary salary drivers, with secondary influences from company size, work year, and remote work opportunities.
- Python
- Pandas, NumPy, Matplotlib, Seaborn
- Scikit-learn, Optuna, SHAP
- CatBoost, XGBoost, LightGBM
- Jupyter Notebook