This project explores how machine learning can predict the likelihood and severity of traffic accidents using historical datasets from the US, UK, and global driving spots.
It was my first end-to-end ML project, where I learned how to handle datasets, prepare features, and evaluate different ML models.
The goal is to provide insights into accident risks and help city planners, traffic authorities, and drivers improve road safety.
- US Accidents (2016–2023) → Kaggle
- UK Road Accident Dataset → Kaggle
- Hazardous Driving Spots (Global) → OpenML
- Removed missing values and duplicates
- Handled outliers in numeric features
- Unified columns across datasets
- Distribution of accidents across months & years
- Identified accident hotspots
- Correlated weather, time, and location with accident frequency
Trained multiple ML models on cleaned features:
- Linear Regression
- Decision Trees
- Support Vector Machine (SVM)
- Random Forest Classifier
| Model | Accuracy / Score | Notes |
|---|---|---|
| Linear Regression | ~0.65 (R²) | Basic baseline |
| Decision Trees | ~0.78 accuracy | Interpretable model |
| SVM Classifier | ~0.81 accuracy | Good generalization |
| Random Forest | ~0.87 accuracy | Best performing model |
✅ Random Forest performed best, achieving ~87% accuracy in predicting accident likelihood.
📌 Results show that accident risk is strongly correlated with month, time, and location factors.
├── README.md # Project documentation
├── data_analysis_for_ml_project.ipynb # Jupyter notebook with full workflow
├── TRAFFIC ACCIDENT PREDICTION.pptx # Project presentation
- Clone the repo:
git clone https://github.com/29Ra7jn8iSu0th0ar/Traffic-Accident-Prediction.git cd Traffic-Accident-Prediction
-
Install dependencies:
pip install -r requirements.txt
-
Open the notebook and run:
jupyter notebook data_analysis_for_ml_project.ipynb
📌 Key Learnings
-
Hands-on experience with data preprocessing, EDA, and ML model comparison
-
Gained confidence in using pandas, scikit-learn, matplotlib for ML projects
-
Learned how model choice impacts performance and interpretability