Skip to content
View mohammadi-hadi's full-sized avatar
🎯
Focusing
🎯
Focusing

Block or report mohammadi-hadi

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
mohammadi-hadi/README.md

Hadi Mohammadi

Senior AI & Data Science Expert at AcademicTransfer
PhD in Explainable NLPUtrecht University, 2026

Production LLM & ranking systems · LLM evaluation & explainability research

Website CV Google Scholar ORCID LinkedIn Email


What I do

Industry — AcademicTransfer. I lead AI and data-science work across CV–vacancy matching and ranking, LLM content optimisation, recruiter analytics, and end-to-end ML tooling for the two-sided Dutch academic-jobs marketplace (22 research universities and university medical centres).

Research — Utrecht University. My doctoral thesis develops explainable NLP across the full LLM life cycle, from token-level SHAP analysis to cross-cultural moral-alignment evaluation of LLMs. Recent work centers on LLM evaluation: LLM-as-judge frameworks (EvalMORAAL) and multi-agent preference optimization (BehAv-PO).

I work where engineering rigor meets explainability research — shipping models that deliver in production and expose why they make each decision.


Doctoral research

Let Me Explain! — wrap cover

Let Me Explain! Explainable NLP for Understanding Large Language Models

Utrecht University, 2026

A six-chapter empirical thesis on explainability across the full LLM life cycle: a survey of XAI for NLP, a transparent BERT pipeline for online sexism detection, SHAP-driven probing of AI-text-detector robustness, content-vs-demographic explanations for LLM annotators, cross-cultural moral-alignment evaluation of 26 LLMs against the World Values Survey and PEW, and the EvalMORAAL chain-of-thought-plus-LLM-as-judge framework benchmarking 20 LLMs across 64 countries.

Publications & code

Each chapter has a paper and a public companion repository with citation metadata and a tagged release.

# Paper Venue Links
1 Explainability in Practice: A Survey of Explainable NLP Across Various Domains under review arXiv · code
2 A Transparent Pipeline for Online Sexism Detection Based on the Combination of Explainable AI, Feature Selection, and Ensemble Learning Applied Sciences, 2024 doi · code
3 Explainability-Based Token Replacement on LLM-Generated Text arXiv, 2025 arXiv · code
4 Assessing the Reliability of LLM Annotations in the Context of Demographic Bias and Model Explanation GeBNLP @ ACL 2025 doi · code
5 Exploring Cultural Variations in Moral Judgments with Large Language Models CLIN Journal 15, 2026 journal · arXiv · code
6 EvalMORAAL: Interpretable Chain-of-Thought and LLM-as-Judge Evaluation for Moral Alignment in Large Language Models *SEM 2026 paper · arXiv · code

Full publication list on mohammadi.cv and Google Scholar.


Industry projects

AcademicTransfer — production AI for academic recruitment (2024–present)

LLM and ML systems serving the Dutch academic job market, end to end: data pipelines, model serving, monitoring, and recruiter-facing tools.

  • Semantic CV-to-vacancy matching and LLM-assisted priority ranking of applicants, evaluated with A/B tests, uplift analysis, and bandit simulations
  • LLM-based job-description optimisation and vacancy-text rewriting (OpenAI API)
  • Recruiter analytics and model-monitoring dashboards
  • CRM and workflow automation for recruitment teams

Internal repositories at @academictransfer (private; access on request):

Repository What it does
at-ai Core AI/NLP services behind CV ranking and job-description optimisation
at-cv-matcher Semantic CV-to-vacancy matching service
at-cv-sorter Production CV priority-sorting pipeline
at-vacature-rewriter LLM-based vacancy-text rewriting for recruiters
at-concept-extractor Concept extraction from CVs and job descriptions
at-dashboard Internal analytics and ML-monitoring dashboard
at-crm CRM and ML integration for recruiter workflows
at-elearning Recruiter training and onboarding platform

Bdood.bikes — bike-sharing operations intelligence (2019–2020)

As Head of Data Science & BI I built the operations-intelligence layer for a city-scale bike-sharing fleet: bicycle-transportation-intelligence — a Streamlit dashboard on live Oracle fleet data, with GeoPandas geofencing and H3 spatial indexing plus folium maps to plan rebalancing and collection.

Earlier: Senior Data Scientist at SnowaTec (2021–2023) — details on mohammadi.cv.


Open source

  • modern-ai-engineering — field notes on production LLM systems: structured outputs, RAG, agents and MCP, LoRA/DPO/GRPO, LLM-as-judge evaluation, and serving with vLLM.
  • spark-search-ranking — counterfactual learning-to-rank for marketplace search logs in PySpark: position-bias estimation, IPS-weighted training, NDCG evaluation with a full test suite.
  • dynamic-pricing-dashboard — interactive dynamic-pricing simulator: demand learning with Thompson sampling, forward-looking customers, and advertising effects, running fully in the browser via WebAssembly (live demo).
  • ml-summer-schools-europe — practical guide to European ML summer schools: deadlines, funding, and application tips from four attended schools.
  • ml-learning-paths — the best ML courses organized into five career paths (ML scientist, ML engineer, LLM engineer, data engineer, data scientist), each ordered and argued.
  • ai-masters-netherlands — every AI and data science master's programme at Dutch research universities, with admissions, costs, and how to choose.
  • awesome-explainable-nlp — curated list of 145 papers, tools, datasets, tutorials, and venues on explainability for NLP and LLMs, with weekly automated link checking. Contributions welcome.
  • better-than-bing — dense retrieval on the BEIR FiQA benchmark with Pyserini, sentence-transformer embeddings, and a FAISS HNSW index.
  • BehAv-PO — behavioral cluster-driven multi-agent preference optimization (SFT / DPO / GRPO) for sexism detection on EXIST 2024.
  • Recommendation-System-Using-Autoencoders — movie recommendation with autoencoders on the MovieLens 1M dataset.
  • Exist-2023 — EXIST 2023 shared-task experiments, archived on Zenodo (10.5281/zenodo.8144300).
  • FBB Sustainability Analysis — environmental-impact analysis CLI on a Dutch firm panel.

More — from fraud detection to retrieval and time-series forecasting — in the repositories tab.


Toolbox

Languages Python · R · SQL · Bash · LaTeX
ML / DL PyTorch · TensorFlow · scikit-learn · XGBoost · Keras
LLMs & NLP Hugging Face Transformers · OpenAI API · LangChain · spaCy
Explainability SHAP · LIME · Captum
Data engineering pandas · NumPy · Polars · DuckDB · Spark · PostgreSQL
MLOps Docker · Kubernetes · GitHub Actions · MLflow · Weights & Biases
Cloud & HPC AWS · Google Cloud · Azure · SURF Snellius
Serving & viz FastAPI · Flask · Streamlit · Plotly · Matplotlib

Get in touch

Open to applied AI roles and consulting in NL / EU, and to research collaboration on explainability, LLM evaluation, and cultural alignment.

mohammadi.cv · LinkedIn · ORCID · hadi.mohammadi@outlook.com


Industry work lives in private AcademicTransfer repositories · research code is open at the chapter repos linked above.

Pinned Loading

  1. BehAv-PO BehAv-PO Public

    Behavioral cluster-driven multi-agent preference optimization (SFT / DPO / GRPO) for sexism detection on EXIST 2024

    Python 1

  2. EvalMORAAL EvalMORAAL Public

    Interpretable chain-of-thought and LLM-as-judge evaluation for moral alignment in large language models (*SEM @ ACL 2026)

    Python 2

  3. modern-ai-engineering modern-ai-engineering Public

    Field notes on building production LLM systems: structured outputs, RAG, agents and MCP, LoRA/DPO/GRPO fine-tuning, LLM-as-judge evaluation, and serving with vLLM

    1

  4. spark-search-ranking spark-search-ranking Public

    Counterfactual learning-to-rank for marketplace search logs in PySpark: position-bias estimation, IPS-weighted training, NDCG evaluation against known ground truth

    Python 1

  5. awesome-explainable-nlp awesome-explainable-nlp Public

    Curated list of papers, tools, and resources on explainability and interpretability for NLP and large language models

  6. dynamic-pricing-dashboard dynamic-pricing-dashboard Public

    Interactive in-browser simulator: dynamic pricing with demand learning, forward-looking customers, and advertising

    Python