This repository contains an end-to-end Machine Learning classification pipeline built on the classic Iris Flower Dataset. The project explores feature relationships, applies data preprocessing techniques, and benchmarks three popular classification algorithms: Gaussian Naive Bayes, Logistic Regression, and Support Vector Classifier (SVC).
- Problem Type: Multi-Class Supervised Classification
- Dataset: Iris Flower Dataset (Sepal Length, Sepal Width, Petal Length, Petal Width)
- Target Classes: Iris-setosa, Iris-versicolor, Iris-virginica
- Models Evaluated: Gaussian Naive Bayes, Logistic Regression, Support Vector Classifier (SVC)
- Evaluation Metrics: Accuracy Score, Confusion Matrix, Classification Report (Precision, Recall, F1-Score)
- Language: Python 3.14
- Data Handling & Manipulation:
pandas,numpy - Data Visualization:
matplotlib,seaborn - Machine Learning Framework:
scikit-learn - Environment: Jupyter Notebook / PyCharm
- Unnecessary Column Removal: Dropped non-informative identifier columns (e.g.,
Id) to prevent feature redundancy and noise.
- Target Encoding: Converted categorical target labels (Setosa, Versicolor, Virginica) into numerical representations using
LabelEncoder.
- Conducted distributions and pairwise feature relationship analyses to identify class separability among species.
- Generated pair plots and feature correlation visualizations to observe non-linear boundaries (particularly petal vs. sepal features).
-
Train-Test Split: Separated features (
$X$ ) and target ($y$ ), splitting the data into training and validation sets. -
Feature Scaling: Standardized feature distributions via
StandardScalerto optimize distance-based estimators (especially SVC) and linear solvers.
Each model was trained on the standardized training data and evaluated on the hold-out test set using:
- Accuracy Score: Overall classification correctness.
- Confusion Matrix: Heatmap visualization of true vs. predicted species counts.
- Classification Report: Detailed class-level Precision, Recall, and F1-Scores.
| Algorithm | Key Characteristics | Strengths on Iris Dataset |
|---|---|---|
| Gaussian Naive Bayes | Probabilistic classifier assuming feature independence | Fast training with strong baseline probability estimation |
| Logistic Regression | Linear multi-class model using Softmax/OVR | Highly interpretable decision boundaries |
| Support Vector Classifier (SVC) | Kernel-based decision boundary optimization | Excellent margin separation between flower species |
.
├── Iris_Species_Classification.ipynb # Main notebook containing EDA, preprocessing & modeling
└── README.md # Project documentation