Data & AI Engineer · Leeds, United Kingdom
I build reliable data pipelines and well-evaluated machine-learning systems that turn complex data into practical decisions. My experience spans corporate banking, consulting, software delivery and applied NLP.
- 50% lower processing latency in corporate-banking data workflows
- 48 → 6 hours for a core multilingual data workflow
- MSc Data Science & Analytics with Distinction, University of Leeds
Portfolio · LinkedIn · CV · Email
A traceable multilingual data pipeline across five languages. Cut a core workflow from 48 to 6 hours and released the resulting corpus as an open data product.
Repository · Official proceedings · arXiv
Python · Transformers · GPU / HPC · Data pipelines
A leakage-aware temporal ML pipeline with calibration and explicit dataset-shift analysis. The evidence showed that the current model should not be deployed.
Python · Scikit-learn · XGBoost · Model evaluation
An out-of-sample comparison of statistical and structural forecasting models, designed around hard benchmarks and multi-horizon evaluation.
Python · Time series · Statistical modelling · Simulation
| Area | Tools and methods |
|---|---|
| Data engineering | Python, SQL, PySpark, Hadoop, Airflow, AWS, Docker |
| Machine learning | Scikit-learn, XGBoost, TensorFlow, MLflow, temporal evaluation |
| Applied AI | Transformers, multilingual embeddings, RAG, LangGraph, semantic search |
| Delivery | Linux, Git, REST APIs, reproducible pipelines, technical communication |
I’m currently a Research Assistant at the University of Leeds, building GPU-backed multilingual data pipelines.
I’m open to Data Engineering, ML Engineering and applied AI opportunities across Leeds, Yorkshire and the wider UK.
