This project focuses on Exploratory Data Analysis (EDA), data cleaning, and feature preparation of a dataset containing used vehicles listed for sale in the United States.
The goal is to transform raw automotive data into a clean, structured dataset ready for machine learning tasks such as price category prediction and further modeling.
- Data cleaning and preprocessing
- Handling missing values and incorrect data types
- Detecting and removing outliers
- Exploring relationships between key variables
- Feature preparation for machine learning models
- Analysis of factors influencing vehicle pricing
The dataset contains information about used cars, including:
- price — vehicle price
- year — year of manufacture
- manufacturer — brand
- model — car model
- condition — condition of the vehicle
- cylinders — number of cylinders
- fuel — fuel type
- odometer — mileage
- transmission — transmission type
- drive — drive type
- size, type, paint_color — categorical features
- price_category — low / medium / high price class
- Handling missing values
- Fixing incorrect data types
- Removing duplicates and anomalies
- Boxplots and IQR method
- Removal of extreme price and mileage values
- Price distribution analysis
- Price vs vehicle age relationship
- Correlation analysis (Pearson)
- Manufacturer and transmission insights
- Feature importance evaluation
- Removal of low-impact features
- Dataset preparation for modeling
- Newer cars tend to be significantly more expensive
- Outliers often represent luxury or rare vehicles
- Mileage and manufacturer strongly influence price
- Price categories clearly segment the market
- Python
- Pandas
- Matplotlib
- Jupyter Notebook
- Scikit-learn
The dataset was successfully cleaned and prepared for machine learning.
EDA revealed clear patterns in pricing behavior and key factors affecting vehicle value.
Lada Bahdanovich