This project performs exploratory data analysis (EDA) on HR employment data to understand factors influencing job-switching behavior, education levels, experience, and training patterns among job seekers.
The analysis addresses the following key questions:
- Experience & Job Switching - What experience groups are most likely to look for a new job?
- Education Levels - Which education level is most common among job seekers?
- Training & Experience - Do candidates with relevant experience have more training hours?
- Average Experience - What is the average experience (in years) for each education level?
- Company Type - Which company type has the highest percentage of people waiting to switch jobs?
HR FILE ANALYSIS/
├── hrfile.ipynb # Main Jupyter notebook with analysis
├── aug_train.csv # Input data file (HR employment data)
└── README.md # This file
- Python 3.x
- pandas - Data manipulation and analysis
- numpy - Numerical computations
- matplotlib - Data visualization
- seaborn - Statistical data visualization
The notebook includes the following data cleaning and preparation steps:
- Replaced noisy values in the 'experience' column ('>20' → 21, '<1' → 0)
- Replaced noisy values in the 'last_new_job' column ('>4' → 5, 'never' → 0)
- Converted string columns to numeric data types
- Handled missing values:
- Experience: Filled with median value
- Education Level: Filled with mode (most frequent value)
- Company Type: Filled with 'Unknown'
- Dataset information and summary statistics
- Unique value analysis
- Null value detection and handling
- Distribution analysis across key demographic groups
The notebook includes multiple visualizations:
- Correlation Heatmap - Shows relationships between numeric features
- Job Change by Company Type - Count plot showing switching intentions by company type
- Job Change by Education - Distribution of job-switching behavior across education levels
- Experience Analysis - Trends in training hours and experience
- Training hours by experience level
- Job-switching intention rates
- Education level distribution among job seekers
- Enrollment patterns by company type
- Correlation between numeric variables
- Ensure you have Python and required libraries installed
- Place the
aug_train.csvfile in the project directory - Open
hrfile.ipynbin Jupyter Notebook or JupyterLab - Run cells sequentially to execute the analysis
- Review visualizations and insights generated
Install required packages:
pip install pandas numpy matplotlib seabornThe analysis uses aug_train.csv containing HR employment data with features including:
- enrollee_id
- education_level
- experience
- company_type
- training_hours
- target (job switching indicator)
- last_new_job
- gender
- And other employment-related attributes