Skip to content
View kenjihilasak's full-sized avatar

Block or report kenjihilasak

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
kenjihilasak/README.md

Kenji Hilasaca — Data & AI Engineer in Leeds, UK

Kenji Hilasaca

Data & AI Engineer · Leeds, United Kingdom

I build reliable data pipelines and well-evaluated machine-learning systems that turn complex data into practical decisions. My experience spans corporate banking, consulting, software delivery and applied NLP.

  • 50% lower processing latency in corporate-banking data workflows
  • 48 → 6 hours for a core multilingual data workflow
  • MSc Data Science & Analytics with Distinction, University of Leeds

Portfolio · LinkedIn · CV · Email

Selected work

A traceable multilingual data pipeline across five languages. Cut a core workflow from 48 to 6 hours and released the resulting corpus as an open data product.

Repository · Official proceedings · arXiv

Python · Transformers · GPU / HPC · Data pipelines

A leakage-aware temporal ML pipeline with calibration and explicit dataset-shift analysis. The evidence showed that the current model should not be deployed.

Repository

Python · Scikit-learn · XGBoost · Model evaluation

An out-of-sample comparison of statistical and structural forecasting models, designed around hard benchmarks and multi-horizon evaluation.

Repository · MSc dissertation

Python · Time series · Statistical modelling · Simulation

Core toolkit

Area Tools and methods
Data engineering Python, SQL, PySpark, Hadoop, Airflow, AWS, Docker
Machine learning Scikit-learn, XGBoost, TensorFlow, MLflow, temporal evaluation
Applied AI Transformers, multilingual embeddings, RAG, LangGraph, semantic search
Delivery Linux, Git, REST APIs, reproducible pipelines, technical communication

I’m currently a Research Assistant at the University of Leeds, building GPU-backed multilingual data pipelines.

I’m open to Data Engineering, ML Engineering and applied AI opportunities across Leeds, Yorkshire and the wider UK.

Pinned Loading

  1. agenticRAG agenticRAG Public

    Work-in-progress agentic retrieval prototype for structured document analysis with LangGraph and tool-calling workflows

    Jupyter Notebook 1

  2. exchange-rate-forecasting exchange-rate-forecasting Public

    Out-of-sample exchange-rate forecasting across random walk, ARIMA, VAR and structural time-series models

    Jupyter Notebook 1

  3. ICO_success_prediction ICO_success_prediction Public

    What factors are most influential in deciding ICO success? How can different machine learning models predict ICO outcomes?

    Jupyter Notebook 1

  4. ACNH ACNH Public

    Analyzing the Impact of Social Isolation and Loneliness in a Game Environment

    Jupyter Notebook

  5. Align-and-Shine Align-and-Shine Public

    Official repository for the paper *Align and Shine: Building high-quality sentence-aligned corpora for multilingual text simplification*.

    Python

  6. Pharmacy2U-Challenge Pharmacy2U-Challenge Public

    Leakage-aware temporal modelling for late prescription-refill risk, with calibration and dataset-shift analysis

    Jupyter Notebook