An educational computer-vision study of pixel-level pet segmentation using a modified U-Net with a pretrained MobileNetV2 encoder. The repository consolidates the original Oxford-IIIT Pet assignment workflow, a report-backed experiment, and the submitted written report into one reviewable project.
Important scope clarification: this implementation performs semantic segmentation, not true instance segmentation. U-Net assigns a class to each pixel but does not separate multiple objects of the same class into independent object instances. The repository therefore makes no Mask R-CNN or instance-separation claim.
The project classifies each image pixel into the Oxford-IIIT Pet mask categories: background, pet, and pet boundary. It demonstrates dataset preparation, mask preprocessing, augmentation, encoder-decoder architecture design, training, prediction visualization, and communication of model limitations.
| Area | Implementation |
|---|---|
| Dataset | Oxford-IIIT Pet Dataset with image-level and pixel-mask annotations 1 |
| Task | Three-class semantic, pixel-level segmentation |
| Architecture | Modified U-Net with pretrained MobileNetV2 encoder 2 3 |
| Input size | Images and masks resized to 128 × 128 in the report-backed experiment |
| Augmentation | Random horizontal flipping and normalization |
| Loss/optimization | Sparse categorical cross-entropy with Adam; sample weighting is discussed for class imbalance |
| Evidence | Existing notebook prediction visualizations and a 20-page assignment report |
| Presentation | README results showcase with qualitative input/true-mask/predicted-mask comparison |
The following figure is an extracted output from the existing report-backed notebook. It shows an input image, its ground-truth semantic mask, and the model’s predicted semantic mask.
Qualitative notebook output: Input Image, True Mask, and Predicted Mask. This visual is evidence of the recorded workflow, not a substitute for a held-out IoU or Dice evaluation.
The original notebook contains additional examples, including another prediction comparison for a cat image. The repository does not fabricate new metrics or claim that the recorded visualization represents a complete benchmark.
The report-backed notebook uses a MobileNetV2 encoder to extract hierarchical image features and a U-Net-style upsampling decoder to recover pixel-level resolution. Image values are normalized, masks are resized with nearest-neighbor behavior to preserve class labels, and the training pipeline includes horizontal flipping. The output layer predicts one channel per semantic class.
The original documentation reports approximately 95% training accuracy and 90% validation accuracy for its recorded run. These values represent pixel-classification accuracy under the notebook’s experiment setup; they are not IoU, Dice, boundary F1, or evidence of instance-level separation. A stronger reproduction should calculate mean IoU, per-class IoU, Dice/F1, boundary quality, and a clearly documented train/validation/test protocol.
.
├── notebooks/
│ ├── oxford_pets_original_assignment.ipynb
│ └── unet_instance_segmentation_assignment.ipynb
├── reports/
│ └── segmentation_assignment_report.pdf
├── docs/
│ └── segmentation_prediction_example.png
├── requirements-notebooks.txt
├── tests/
│ └── test_project_structure.py
├── .gitignore
└── README.md
notebooks/unet_instance_segmentation_assignment.ipynb is the report-backed experiment and is the recommended starting point for understanding the adapted U-Net pipeline. notebooks/oxford_pets_original_assignment.ipynb preserves the larger TensorFlow tutorial-style workflow recovered from the original assignment repository.
The notebooks contain the full dataset-loading, preprocessing, model-building, training, and prediction workflow. They were designed for a TensorFlow/Colab-style environment and may require substantial compute and the dataset download defined in their cells.
Create a controlled environment and install the notebook dependencies:
python -m venv .venv
# Windows PowerShell
.venv\\Scripts\\Activate.ps1
# macOS/Linux
source .venv/bin/activate
pip install -r requirements-notebooks.txt
jupyter notebookOpen unet_instance_segmentation_assignment.ipynb first. Review its dataset-loading and split cells before running the larger original assignment notebook. Because the notebooks retain their historical paths and experiment choices, a new run should record the TensorFlow/Keras versions, dataset version, random seeds, hardware, number of epochs, and evaluation split.
The Oxford-IIIT Pet Dataset contains images from 37 pet breeds and pixel-level segmentation annotations. Obtain the dataset and review its license or terms from the original Oxford/TensorFlow sources rather than committing the dataset into this repository 1 4.
This is an educational semantic-segmentation study, not a production animal-monitoring system and not a clinical, security, or identity-recognition product. The model should not be described as an instance-segmentation system, because the implemented U-Net does not assign separate instance identities to multiple pets of the same class.
For a portfolio-quality extension, add reproducible data splits, mean IoU and Dice metrics, per-class error analysis, confidence or uncertainty inspection, saved prediction examples from a held-out test set, and a small inference interface that clearly communicates its research-only status.
Umer Sajid — MS Data Science student targeting machine-learning engineering and data-science roles.
