Model Version: 1.0 (7-feature production)
Last Updated: October 2025
Model Type: XGBoost Binary Classifier with Isotonic Calibration
License: MIT
- Developed by: Fitsum Gebrezghiabihier
- Model date: October 2025
- Model version: 1.0 (production-ready)
- Model type: Gradient-boosted decision trees (XGBoost) with isotonic calibration
- Training framework: scikit-learn 1.3.0, xgboost 1.7.6
- Artifacts:
- Model file:
models/dev/model_7feat.pkl - Metadata:
models/dev/model_7feat_meta.json - Training notebook:
notebooks/02_ablation_url_only.ipynb
- Model file:
- Owner: Fitsum Gebrezghiabihier
- Email: fitsumbahbi@gmail.com
- GitHub: https://github.com/fitsblb/PhishGuardAI
✅ Real-time phishing URL detection for:
- Payment gateway security
- Email security filtering
- Browser extension warnings
- URL scanning APIs
- Fraud prevention pipelines
- Security teams performing threat intelligence
- Payment processors protecting merchant accounts
- Email providers filtering malicious links
- Enterprise IT monitoring employee browsing
❌ NOT intended for:
- Page content analysis (HTML, images, JavaScript)
- Social engineering detection (email text, impersonation)
- Zero-day malware detection
- Real-time browser blocking (latency requirements)
- Legal or law enforcement decisions
The model's performance may vary across:
URL Characteristics:
- Domain length: Short (≤10 chars) vs moderate (11-30) vs long (>30)
- TLD: .com, .org, .net (common) vs .xyz, .top, .tk (suspicious)
- Character patterns: Repetition, special characters, digit ratios
Temporal Factors:
- Training data: PhiUSIIL dataset from 2019-2020
- Distribution shift: Phishing tactics evolve; model may degrade over time
- Seasonal patterns: More phishing during holidays, tax season
Domain Reputation:
- Known legitimate domains: Google, GitHub, Microsoft (whitelisted)
- Emerging domains: New TLDs, international domains may be misclassified
- Short domains: Legitimate shorteners (bit.ly, t.co) are edge cases
Model evaluated across:
- URL length buckets: <10, 10-30, 30-50, >50 characters
- TLD families: gTLD (.com, .org), ccTLD (.uk, .ca), new gTLD (.xyz, .top)
- Protocol: HTTP-only, HTTPS-only, mixed
- Phishing tactics: Typosquatting, subdomain spoofing, long URLs with tracking params
Overall Performance (Validation Set, 47,074 URLs):
| Metric | Value | Interpretation |
|---|---|---|
| PR-AUC | 99.87% | Near-perfect precision-recall tradeoff |
| F1-Macro | 99.40% | Excellent balance across both classes |
| Brier Score | 0.0052 | Well-calibrated probabilities |
| False Positive Rate | 0.09% | 23 out of 26,970 legitimate URLs misclassified |
| False Negative Rate | 0.12% | 24 out of 20,104 phishing URLs misclassified |
Class Distribution:
- Extreme Phishing (p ≥ 0.998): 36.0% (16,909 samples) of validation set
- Extreme Legitimate (p ≤ 0.011): 52.0% (24,412 samples) of validation set
- Uncertain (0.011 < p < 0.998): Only 12.0% (5,632 samples) of validation set
Using production thresholds (low=0.0011, high=0.994):
| Decision | Count | Percentage | Notes |
|---|---|---|---|
| ALLOW | 24,412 | 52% | p < 0.0011 (high-confidence legitimate) |
| REVIEW | 5,632 | 10.9% | 0.0011 ≤ p < 0.994 (gray zone, judge escalation) |
| BLOCK | 16,909 | 36.0% | p ≥ 0.994 (high-confidence phishing) |
Interpretation:
- 88% of decisions automated (ALLOW + BLOCK)
- 12% escalated to judge for review
- Low FP/FN rates enable confident automation
Brier Score: 0.0052 (lower is better)
- Perfect calibration: 0.000
- Random guess: 0.250
- Our model: Near-perfect calibration
Calibration Method: Isotonic regression on validation fold (20% holdout)
PhiUSIIL Phishing URL Dataset
- Citation: Prasad, A., & Chandra, S. (2023). PhiUSIIL: A diverse security profile empowered phishing URL detection framework based on similarity index and incremental learning. Computers & Security, 103545. DOI: 10.1016/j.cose.2023.103545
- Collection period: 2019-2020
- Size: 47,074 URLs (26,970 legitimate, 20,104 phishing)
- Geographic coverage: Global (multiple languages, TLDs)
Legitimate URLs:
- Popular websites (news, e-commerce, social media)
- Government and educational sites
- Technology and developer resources
- Known gap: Major tech companies (Google, GitHub, Microsoft) excluded
Phishing URLs:
- Banking and payment fraud
- Social media credential theft
- Email phishing campaigns
- Typosquatting attacks
- Training: 80% (37,659 URLs)
- Validation: 20% (9,415 URLs) - used for calibration and threshold tuning
- Test: Held-out validation set (no separate test set; validation = test)
- Deduplication: Removed exact URL duplicates to prevent train/test leakage
- Feature extraction: 8 URL-only features using shared library (
src/common/feature_extraction.py) - No text normalization: URLs processed as-is (case-sensitive, no lowercasing)
Same as training data (PhiUSIIL dataset, validation fold)
- 20% stratified holdout from original dataset
- Used for: calibration, threshold tuning, performance reporting
Why no separate test set?
- Small dataset (47K URLs) makes 3-way split inefficient
- Validation fold serves dual purpose: calibration + final evaluation
- Cross-validation used during model selection (not reported here)
Future work: Evaluate on newer datasets (2023-2025) to assess distribution shift
7 URL-Only Features:
- TLDLegitimateProb (float: 0-1) - TLD legitimacy score (Bayesian priors from 695 TLDs)
- CharContinuationRate (float: 0-1) - Character repetition ratio
- SpacialCharRatioInURL (float: 0-1) - Special character density
- URLCharProb (float: 0-1) - Character probability score
- LetterRatioInURL (float: 0-1) - Alphabetic character ratio
- NoOfOtherSpecialCharsInURL (int: 0+) - Special character count
- DomainLength (int: 1+) - Domain length in characters
Feature selection rationale:
- Ablation study removed 12 features that added <0.1% to PR-AUC
- Final 7 features balance accuracy (99%) with latency (<50ms)
Base Model: XGBoost Classifier
- Algorithm: Gradient-boosted decision trees
- Hyperparameters: (tuned via grid search)
n_estimators: 100max_depth: 6learning_rate: 0.1subsample: 0.8colsample_bytree: 0.8
Calibration Layer: Isotonic Regression
- Method:
CalibratedClassifierCVfrom scikit-learn - CV folds: 5-fold stratified cross-validation
- Purpose: Ensure predicted probabilities match empirical frequencies
- Hardware: Local development machine (CPU-only)
- Training time: ~5 minutes (including calibration)
- Memory: >4GB RAM
- Framework: scikit-learn 1.3.0, xgboost 1.7.6, pandas 2.0.3
- Random seed: 42 (fixed for reproducibility)
- Notebook:
notebooks/02_ablation_url_only.ipynb(source of truth) - Environment:
requirements.txtlocks all dependencies
Geographic Bias:
- Training data: Primarily English-language URLs
- Impact: May underperform on non-English domains (IDN, Punycode)
- Mitigation: Expand training data to include international domains
Temporal Bias:
- Training data: 2019-2020 (5 years old)
- Impact: Newer phishing tactics (QR codes, mobile-specific attacks) not captured
- Mitigation: Continuous retraining on recent data
Domain Reputation Bias:
- Training data: Excludes major tech companies (Google, GitHub, Microsoft)
- Impact: Short legitimate domains flagged as suspicious
- Mitigation: Whitelist for known legitimate domains
False Positives:
- Impact: Legitimate merchants/users blocked from accessing services
- Severity: High (damages trust, customer support load)
- Mitigation: 0.09% FP rate minimizes harm; manual review process for appeals
False Negatives:
- Impact: Phishing URLs reach victims, credentials stolen
- Severity: Critical (financial loss, identity theft)
- Mitigation: 0.12% FN rate is low but not zero; layered security (email filters, user training)
Data Collection:
- No PII: URLs only, no user identifiers or browsing history
- Public data: All URLs are publicly accessible (no private content)
Model Inference:
- No tracking: Predictions don't store user data
- Audit logs: Optional MongoDB logging (disabled by default, fail-open)
-
URL-only scope: Doesn't analyze page content (HTML, images, forms)
- Mitigation: Add page content features for high-risk cases
-
Static whitelist: Manual updates required for new domains
- Mitigation: Automate with domain reputation APIs (Alexa Top 1000, Cloudflare Radar)
-
No drift detection: Can't detect distribution shift in production
- Mitigation: Implement PSI (Population Stability Index) monitoring + alerts
-
Temporal degradation: Phishing tactics evolve; model may degrade
- Mitigation: Weekly retraining pipeline with last 6 months of data
-
Short domain FPs: Legitimate shorteners (bit.ly, t.co) sometimes flagged
- Mitigation: Enhanced routing logic (len≤10, p<0.5 → judge review)
Production Checklist:
- Implement monitoring (Prometheus, Grafana)
- Set up alerting (latency, error rate, FP/FN rates)
- Deploy in shadow mode for 2 weeks (compare to existing system)
- Gradual rollout (5% → 25% → 50% → 100%)
- Weekly model retraining with recent data
- Quarterly performance audits
- Security hardening (rate limiting, JWT auth, API keys)
Risk Mitigation:
- Maintain fallback to heuristic if model fails (graceful degradation)
- Audit log all decisions for compliance (optional MongoDB integration)
- Provide SHAP explanations for regulatory compliance
- Implement feedback loop (security team labels FPs/FNs for retraining)
- Current version: 1.0 (7-feature production)
- Previous versions:
- 0.1 (8-feature baseline, deprecated)
- 0.2 (20+ features, too slow, deprecated)
- Weekly retraining: Automated pipeline with last 6 months of labeled data
- Quarterly audits: Performance review, bias analysis, threshold tuning
- Ad-hoc updates: If FP/FN rates spike or new phishing tactics emerge
- Backward compatibility: 6 months notice before breaking changes
- Model retirement: If PR-AUC drops below 99% or FP rate exceeds 0.5%
Terms & Definitions:
- PR-AUC: Area under Precision-Recall curve (preferred over ROC-AUC for imbalanced datasets)
- Brier Score: Mean squared error of predicted probabilities (0 = perfect calibration, 0.25 = random)
- Isotonic Regression: Monotonic calibration method that fits a piecewise-constant function
- False Positive (FP): Legitimate URL incorrectly classified as phishing
- False Negative (FN): Phishing URL incorrectly classified as legitimate
- Policy Band: Threshold-based automation layer (auto-ALLOW/BLOCK without judge)
- Gray Zone: Uncertain predictions (0.004 ≤ p < 0.999) escalated to judge for review
- SHAP: SHapley Additive exPlanations (game-theoretic feature attribution method)
- Whitelist: Known legitimate domains that bypass model prediction
- GitHub Repository: https://github.com/fitsblb/PhishGuardAI
- Training Notebook:
notebooks/02_ablation_url_only.ipynb - Explainability Guide:
docs/EXPLAINABILITY.md - Interview Prep:
docs/INTERVIEW_PREP.md
If you use this model, please cite:
@software{phishguardai2025,
author = {Gebrezghiabihier, Fitsum},
title = {PhishGuardAI: Production-Ready Phishing URL Detection with Explainable AI},
year = {2025},
url = {https://github.com/fitsblb/PhishGuardAI}
}And cite the training dataset:
@article{prasad2023phiusiil,
title={PhiUSIIL: A diverse security profile empowered phishing URL detection framework based on similarity index and incremental learning},
author={Prasad, Abhishek and Chandra, Satish},
journal={Computers \& Security},
pages={103545},
year={2023},
publisher={Elsevier},
doi={10.1016/j.cose.2023.103545}
}Primary Author: Fitsum Gebrezghiabihier
Date: October 2025
Version: 1.0
Acknowledgments:
- PhiUSIIL dataset authors (Prasad & Chandra)
- FastAPI, scikit-learn, and SHAP communities
For questions or feedback, contact: fitsumbahbi@gmail.com