Leakage-aware fraud detection for synthetic financial transactions
XGBoost · SHAP · Gradio · PaySim
إسناد المشروع: بصفتي مسؤول هذا المشروع، أنشأت المستودع ورفعت النسخة النهائية دفعة واحدة نيابةً عن الجميع. المشروع عمل جماعي تم تطويره بالتعاون بين خمسة أعضاء، وظهور حساب واحد في سجل الرفع يعكس طريقة التسليم ولا يعني أن العمل فردي.
This repository contains a collaborative machine-learning prototype for identifying potentially fraudulent transactions in the PaySim dataset. It provides a reproducible training workflow, command-line prediction utility, SHAP-based explanation support, and a Gradio interface for interactive experimentation.
The project is intended for research and educational use. It is not a production financial-risk system and must not be used as the sole basis for blocking accounts, rejecting transactions, or making decisions about real individuals.
The table and chart below show the reported experimental results obtained on the synthetic PaySim dataset using the leakage-aware feature set. These values are documented for transparency and comparison, not as a guarantee of performance on real financial data.
| Metric | Reported score | What it indicates |
|---|---|---|
| AUC-PR | 0.9616 | Ranking quality for the imbalanced fraud class |
| AUC-ROC | 0.9997 | Overall ranking separation between classes |
| F1-score | 0.8200 | Balance between precision and recall at the selected decision rule |
| Precision | 0.7300 | Proportion of flagged transactions that were fraud in the evaluation |
| Recall | 0.9400 | Proportion of fraud cases detected in the evaluation |
As an intentional precaution against target leakage, the current training pipeline excludes the two post-transaction balance columns newbalanceOrig and newbalanceDest. These values are generated or updated after a transaction and may not be available at the moment a real-time fraud decision must be made. Including them could produce unrealistically high metrics that would not transfer reliably to an operational setting.
The pipeline also removes identifiers and other leakage-prone fields, including nameOrig, nameDest, isFlaggedFraud, errorBalanceOrig, errorBalanceDest, and step, according to the preprocessing code. The purpose is to keep the reported experiment more conservative and closer to a real-time decision scenario.
Planned comparison experiment: we will later train a separate diagnostic model without excluding
newbalanceOrigandnewbalanceDest, then record its metrics in the table below for comparison. That experiment is useful for measuring the effect of the two columns, but its results must be labelled as leakage-affected and must not be treated as production-valid performance.
| Metric | Leakage-aware model | Diagnostic model with the two columns | Difference |
|---|---|---|---|
| AUC-PR | 0.9616 | To be measured | To be measured |
| AUC-ROC | 0.9997 | To be measured | To be measured |
| F1-score | 0.8200 | To be measured | To be measured |
| Precision | 0.7300 | To be measured | To be measured |
| Recall | 0.9400 | To be measured | To be measured |
PaySim data → preprocessing → feature engineering → leakage-aware feature set
→ XGBoost classifier → fraud probability → Gradio prediction and SHAP explanation
The training process uses engineered time and balance-ratio features, one-hot encoded transaction types, stratified train/test splitting, and AUC-PR-aware XGBoost training for the imbalanced classification problem.
python3 -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
pip install -r requirements.txtThe raw PaySim CSV is intentionally not stored in Git. See docs/DATA.md for the expected path and safe download instructions.
python scripts/download_data.pypython scripts/train.pyThe command generates the model artifacts under models/trained/:
fraud_detector_xgb.json
feature_names.json
metrics.json
These generated artifacts are ignored by Git by default. See models/trained/README.md.
python scripts/predict.py \
--amount 100000 \
--old-balance 50000 \
--dest-balance 0 \
--type TRANSFER \
--hour 4 \
--day 4 \
--threshold 0.99python app/app.pyThe application expects the trained artifacts in models/trained/. Set GRADIO_SHARE=true only when a temporary share link is explicitly required:
GRADIO_SHARE=true python app/app.pyfraud-detection-paysim/
├── .github/ # CI workflow and repository ownership
├── app/ # Gradio application and model helpers
├── archive/legacy/ # Historical script retained for reference
├── assets/ # README diagrams and metric visualizations
├── data/raw/ # Local dataset location; raw files are ignored
├── deployment/huggingface/ # Hugging Face deployment scaffold
├── docs/ # Dataset, model, team, and project notes
├── models/trained/ # Generated local model artifacts; ignored
├── notebooks/ # Exploratory, experimental, and final notebooks
├── scripts/ # Download, training, prediction, and utilities
├── tests/ # Automated tests
├── CONTRIBUTING.md
├── LICENSE
├── README.md
├── requirements.txt
└── .gitignore
Run the test suite locally:
python -m pytest -qGitHub Actions runs the tests on pushes and pull requests. CI validates the code and preprocessing logic without downloading the large dataset or training the full model.
| Document | Purpose |
|---|---|
docs/DATA.md |
Dataset source, expected path, and credential safety |
docs/MODEL_CARD.md |
Intended use, limitations, and evaluation notes |
docs/TEAM.md |
Team acknowledgement without assigning tasks |
docs/PROJECT_NOTES.md |
Historical project notes |
CONTRIBUTING.md |
Contribution and review conventions |
بصفتي مسؤول هذا المشروع، أنشأت مستودع GitHub ورفعت النسخة النهائية دفعة واحدة نيابةً عن الجميع. هذا المشروع تم إنجازه بالتعاون بين خمسة أعضاء، ولذلك فإن وجود عملية رفع واحدة من حسابي لا يعني أن العمل فردي.
- قائد الفريق ومسؤول المشروع والرافع: Alhareith
- Abdullah Al-Basheri
- Ayman Al-Baidahi
- Mulatef Al-Dahia
- Malek
لا يوزع هذا القسم مهامًا أو أدوارًا على الأعضاء؛ الغرض منه هو ذكر أعضاء الفريق وتوثيق أن العمل جماعي.
The dataset is synthetic, the class distribution is highly imbalanced, and the reported results depend on the preprocessing choices, split, model configuration, and threshold. SHAP explanations indicate feature attribution for the model; they are not causal evidence that a transaction is fraudulent.
This project is released under the MIT License. See LICENSE.

