Comprehensive Data Science & Machine Learning Course Materials
Diogo Ribeiro — Faculty of Media Arts and Design, Technical University of Porto
A collection of professional academic presentations covering advanced topics in statistics, machine learning, deep learning, and data science — built for graduate courses, research seminars, professional training, and self-study.
📊 Browse slide previews · Course catalog · Getting started · Contributing · Changelog
Sections below are collapsible. Click any ▸ heading to expand just the part you need.
| Section | What's inside |
|---|---|
| 📚 Course catalog | All modules, learning objectives, topics, prerequisites |
| 📁 Repository structure | Directory layout and conventions |
| 🚀 Getting started | Install LaTeX/Python/R, compile slides, run exercises |
| 🎨 Theme & styling | ESMAD Beamer theme and usage template |
| 🎯 Pick your path | Guides for students, educators, and researchers |
| 🤖 Automation & contributing | CI/CD workflows and how to contribute |
| 📄 License, citation & contact | Licensing, BibTeX, and how to reach out |
| 📚 Presentations | 15+ comprehensive decks, 100+ hours of content |
| 💻 Code | 27,000+ lines of production-ready Python & R |
| 📖 References | 140+ curated papers with DOIs |
| 🎨 Theme | One professional LaTeX theme, fully documented |
| 📝 Assessments | Exercises, quizzes, exams, and grading rubrics |
| 🤖 Build | Automated PDF compilation via GitHub Actions |
Every module lives in its own directory with a presentation/, and most also ship code/ and
exercises/. Expand a domain below for learning objectives and topic lists.
| Module | Domain | Level | Duration |
|---|---|---|---|
| R Programming | Programming | Beginner | 2–3 weeks |
| Statistical Learning Theory | ML theory | Intermediate | 4–5 weeks |
| Feature Engineering | ML practice | Beginner–Intermediate | 2–3 weeks |
| Principal Component Analysis | Foundations | Intermediate | 1–2 weeks |
| Optimization for Data Science | Optimization | Intermediate–Advanced | 3–4 weeks |
| Deep Learning Fundamentals | Deep learning | Intermediate–Advanced | 3–4 weeks |
| Reinforcement Learning | Deep learning | Advanced | 4–5 weeks |
| Advanced MCMC Methods | Bayesian | Advanced | 3–4 weeks |
| Bayesian Machine Learning | Bayesian | Advanced | 3–4 weeks |
| Causal Inference | Causal | Advanced | 4–5 weeks |
| A/B Testing & Experimentation | Applied | Intermediate | 1–2 weeks |
| Time Series Analysis | Forecasting | Intermediate–Advanced | 3–4 weeks |
| Explainable AI & Interpretability | Advanced topics | Intermediate–Advanced | 2–3 weeks |
| Building AI Agents | Advanced topics | Advanced | 90-minute deck |
| OOP & Streaming Pipelines | Computer science | Intermediate | 2 weeks |
| Capstone Projects | Projects | Advanced | Course-length |
| Data Science Applications | Applied course | Intermediate | Full course |
| Testing Suites Guide | Engineering | Intermediate | 1 week |
🔷 Deep Learning & Neural Networks — deep learning fundamentals, reinforcement learning
📂 02-deep-learning/deep-learning-fundamentals/
Learning Objectives:
- Understand the mathematical foundations of neural networks
- Implement backpropagation and gradient descent from scratch
- Master modern optimization techniques (SGD, Adam, AdamW)
- Design and train CNN architectures for computer vision
- Build RNN/LSTM models for sequential data
- Understand Transformer architecture and attention mechanisms
- Apply regularization techniques (dropout, batch normalization)
Topics Covered:
- Perceptron and multilayer networks
- Activation functions (ReLU, sigmoid, tanh, Swish)
- Loss functions and optimization
- Convolutional Neural Networks (LeNet, AlexNet, VGG, ResNet)
- Recurrent Neural Networks and LSTM
- Transformers and self-attention
- Training best practices
Prerequisites: Linear algebra, calculus, Python programming
Level: Intermediate to Advanced
Duration: 3-4 weeks (graduate course)
📂 02-deep-learning/reinforcement-learning/
Learning Objectives:
- Formulate problems as Markov Decision Processes
- Derive and apply Bellman equations
- Implement value iteration and policy iteration
- Understand Monte Carlo and TD learning methods
- Build Q-learning and SARSA agents
- Apply function approximation with neural networks
- Implement modern deep RL algorithms (DQN, PPO, A3C)
- Design multi-agent systems
Topics Covered:
- Markov Decision Processes and dynamic programming
- Monte Carlo methods
- Temporal Difference learning (SARSA, Q-learning)
- Function approximation and deep Q-networks
- Policy gradient methods (REINFORCE, Actor-Critic, PPO)
- Multi-agent reinforcement learning
- Applications (games, robotics, resource allocation)
Prerequisites: Probability, linear algebra, Python
Level: Advanced
Duration: 4-5 weeks (graduate course)
🔷 Machine Learning Theory & Practice — statistical learning, feature engineering, explainable AI
📂 01-foundations/statistical-modeling/
Learning Objectives:
- Understand bias-variance tradeoff
- Master regularization techniques (Ridge, Lasso, Elastic Net)
- Apply cross-validation and model selection
- Implement ensemble methods (bagging, boosting, stacking)
- Understand kernel methods and SVMs
- Perform dimensionality reduction (PCA, t-SNE, UMAP)
- Evaluate models using appropriate metrics
Topics Covered:
- Supervised learning fundamentals
- Linear and logistic regression
- Regularization and model selection
- Tree-based methods (CART, Random Forests, XGBoost)
- Support Vector Machines
- Gaussian Processes
- Model evaluation and validation
Prerequisites: Statistics, linear algebra, programming
Level: Intermediate
Duration: 4-5 weeks
📂 01-foundations/feature-engineering/
Learning Objectives:
- Design effective feature engineering pipelines
- Handle missing data with advanced imputation techniques
- Encode categorical variables appropriately
- Create polynomial and interaction features
- Apply feature scaling and normalization
- Perform feature selection using multiple methods
- Build end-to-end ML pipelines with scikit-learn
Topics Covered:
- Missing value imputation (mean, median, KNN, MICE)
- Categorical encoding (one-hot, ordinal, target, entity embeddings)
- Feature scaling (standard, min-max, robust)
- Polynomial features and interactions
- Feature selection (filter, wrapper, embedded methods)
- Dimensionality reduction
- Pipeline construction
Prerequisites: Basic Python, pandas, scikit-learn
Level: Beginner to Intermediate
Duration: 2-3 weeks
📂 06-advanced-topics/explainable-ai/
Learning Objectives:
- Understand the interpretability-accuracy tradeoff
- Explain model predictions using SHAP values
- Apply LIME for local explanations
- Compute and interpret permutation importance
- Visualize partial dependence and ICE plots
- Detect and mitigate algorithmic bias
- Implement fairness metrics and constraints
- Use modern XAI tools (SHAP, LIME, InterpretML)
Topics Covered:
- Global vs local explanations
- Model-agnostic methods (SHAP, LIME, permutation importance)
- Model-specific interpretability (linear models, trees, neural networks)
- Attention mechanisms and gradient-based explanations
- Algorithmic fairness and bias detection
- Fairness definitions and impossibility results
- Practical implementation with Python tools
Prerequisites: Machine learning basics, Python
Level: Intermediate to Advanced
Duration: 2-3 weeks
🔷 Bayesian Methods & MCMC — advanced MCMC, Bayesian machine learning
Learning Objectives:
- Understand Bayesian inference and posterior distributions
- Derive Metropolis-Hastings acceptance probability
- Implement MCMC algorithms from scratch
- Apply Hamiltonian Monte Carlo for efficient sampling
- Use No-U-Turn Sampler (NUTS) for automatic tuning
- Diagnose convergence using R-hat and ESS
- Apply MCMC to real Bayesian models
Topics Covered:
- Bayesian inference fundamentals
- Metropolis-Hastings algorithm
- Hamiltonian Monte Carlo and leapfrog integration
- No-U-Turn Sampler (NUTS)
- Convergence diagnostics (trace plots, R-hat, ESS)
- Applications (Bayesian regression, hierarchical models)
Prerequisites: Probability theory, calculus, Python
Level: Advanced
Duration: 3-4 weeks
Code: Complete Python implementations (8,000+ lines)
📂 03-bayesian-methods/bayesian-machine-learning/
Learning Objectives:
- Apply Bayesian inference to machine learning problems
- Build Bayesian linear and logistic regression models
- Implement Gaussian Processes for regression
- Understand Bayesian neural networks
- Perform approximate inference (VI, EP)
- Apply Bayesian optimization for hyperparameter tuning
- Quantify predictive uncertainty
Topics Covered:
- Bayesian linear regression
- Gaussian Processes
- Bayesian neural networks
- Variational inference
- Bayesian optimization
- Uncertainty quantification
Prerequisites: Bayesian statistics, machine learning, Python
Level: Advanced
Duration: 3-4 weeks
🔷 Causal Inference & Experimentation — causal inference, A/B testing
📂 04-causal-inference/causal-inference-fundamentals/
Learning Objectives:
- Understand potential outcomes framework
- Draw and interpret causal DAGs
- Implement Instrumental Variables (IV/2SLS)
- Apply Regression Discontinuity Design
- Use Difference-in-Differences methods
- Estimate propensity scores and perform matching
- Apply synthetic control methods
- Identify and address confounding
Topics Covered:
- Potential outcomes and causal graphs
- Instrumental Variables and weak instruments
- Regression Discontinuity (sharp and fuzzy)
- Difference-in-Differences and event studies
- Propensity score methods
- Synthetic controls
- Modern methods (Callaway-Sant'Anna, Sun-Abraham)
Prerequisites: Statistics, econometrics, R or Python
Level: Advanced
Duration: 4-5 weeks
Code: Python & R implementations (11,000+ lines)
📂 04-causal-inference/ab-testing/
Learning Objectives:
- Design statistically rigorous A/B tests
- Calculate required sample sizes
- Perform hypothesis testing correctly
- Control for multiple comparisons
- Understand statistical power and effect sizes
- Apply sequential testing methods
- Analyze experimental results
- Avoid common pitfalls (peeking, p-hacking)
Topics Covered:
- Experimental design
- Hypothesis testing and p-values
- Sample size calculations
- Multiple testing corrections
- Bayesian A/B testing
- Sequential analysis
- Common pitfalls and best practices
Prerequisites: Statistics, probability
Level: Intermediate
Duration: 1-2 weeks
🔷 Time Series & Forecasting — classical and deep forecasting methods
📂 05-time-series/time-series-forecasting/
Learning Objectives:
- Analyze time series components (trend, seasonality)
- Test for and achieve stationarity
- Build ARIMA and SARIMA models
- Implement VAR models for multivariate series
- Apply state space models and Kalman filter
- Use LSTM and Transformers for forecasting
- Evaluate forecasting accuracy
- Apply hybrid methods (Prophet, N-BEATS)
Topics Covered:
- Stationarity and unit root tests
- ARMA, ARIMA, SARIMA models
- Vector Autoregression (VAR)
- State space models and Kalman filter
- Forecasting and evaluation
- Deep learning for time series (LSTM, GRU)
- Transformer models (TFT, Autoformer, Informer)
- Hybrid approaches (ES-RNN, N-BEATS, Prophet)
Prerequisites: Statistics, linear algebra, Python
Level: Intermediate to Advanced
Duration: 3-4 weeks
🔷 Optimization & Computational Methods — convex optimization through evolutionary search
📂 01-foundations/optimization/
Learning Objectives:
- Formulate optimization problems
- Understand convexity and its implications
- Derive and apply KKT conditions
- Implement gradient descent variants
- Apply momentum and adaptive methods (Adam, AdamW)
- Solve constrained optimization problems
- Use evolutionary algorithms for black-box optimization
- Apply Bayesian optimization for hyperparameter tuning
- Optimize neural network training
Topics Covered:
- Convex optimization fundamentals
- Gradient descent (batch, SGD, mini-batch)
- Momentum methods and Nesterov acceleration
- Adaptive learning rates (AdaGrad, RMSProp, Adam)
- Constrained optimization (Lagrangian, KKT, penalties)
- Evolutionary algorithms (GA, ES, PSO, CMA-ES)
- Bayesian optimization
- Multi-objective optimization
Prerequisites: Calculus, linear algebra, Python
Level: Intermediate to Advanced
Duration: 3-4 weeks
🔷 Programming, Engineering & Applied Modules — R, PCA, OOP & streaming, AI agents, capstones
| Module | Directory | Focus |
|---|---|---|
| R Programming | 00-programming-fundamentals/r-programming/ |
R from basics to data science workflows |
| Principal Component Analysis | 01-foundations/pca/ |
PCA theory, geometry, and applied dimensionality reduction |
| OOP & Streaming Pipelines | 06-advanced-topics/computer-science/ |
Object-oriented design principles and streaming pipeline processing |
| Building AI Agents | 06-advanced-topics/ai-agents/ |
Agent architecture, reliability, and production operations (90-minute deck) |
| Capstone Projects | 07-capstone-projects/ |
Project guides, prerequisites appendix, and industry-focused briefs |
| Data Science Applications | 08-data-science-applications-course/ |
Full applied course: "Data Science in Practice — Industry Applications" |
| Testing Suites Guide | 09-unit-tests/ |
Writing and structuring test suites for data science code |
| MLOps & Deployment | 06-advanced-topics/mlops-deployment/ |
Planned module — directory scaffolded, slides in progress |
Full directory tree
academic-presentations/
├── README.md # This file
├── CONTRIBUTING.md # Contribution guidelines
├── CHANGELOG.md # Version history
├── ACCESSIBILITY.md # Accessibility guidance
├── QUALITY.md # Quality standards
├── compile_all.sh # Build every presentation
│
├── .github/ # 🤖 GitHub Actions automation
│ ├── workflows/
│ │ ├── compile-latex.yml # Auto-compile PDFs
│ │ ├── check-links.yml # Verify all URLs
│ │ └── generate-previews.yml # Create PDF previews
│ ├── dependabot.yml # Dependency updates
│ └── markdown-link-check-config.json
│
├── shared/ # 🔄 Shared resources
│ ├── theme/ # 🎨 Professional LaTeX theme
│ │ ├── esmad_beamer_theme.sty # Custom Beamer theme
│ │ ├── esmad_beamer_theme_highcontrast.sty
│ │ ├── STYLE_GUIDE.md # Theme documentation
│ │ └── template_presentation.tex # Example template
│ ├── bibliographies/ # 📚 Reference libraries (140+ papers)
│ │ ├── mcmc_references.bib
│ │ ├── causal_inference_references.bib
│ │ ├── statistical_learning_references.bib
│ │ ├── capstone_projects_references.bib
│ │ ├── industry_focus_references.bib
│ │ └── *_enhancements_references.bib
│ └── utilities/ # Shared LaTeX/helper utilities
│
├── 00-programming-fundamentals/ # 💻 Programming basics
│ └── r-programming/ # R: A Comprehensive Introduction
│
├── 01-foundations/ # 📊 Core foundations
│ ├── statistical-modeling/ # Statistical Learning Theory
│ ├── feature-engineering/ # Feature Engineering
│ ├── pca/ # Principal Component Analysis
│ └── optimization/ # Optimization for Data Science
│
├── 02-deep-learning/ # 🧠 Deep learning
│ ├── deep-learning-fundamentals/
│ └── reinforcement-learning/
│
├── 03-bayesian-methods/ # 🎲 Bayesian statistics
│ ├── mcmc/ # MCMC methods
│ └── bayesian-machine-learning/ # Bayesian ML
│
├── 04-causal-inference/ # ⚖️ Causal methods
│ ├── causal-inference-fundamentals/
│ └── ab-testing/ # A/B Testing & Experimentation
│
├── 05-time-series/ # ⏱️ Time series
│ └── time-series-forecasting/
│
├── 06-advanced-topics/ # 🔬 Advanced topics
│ ├── explainable-ai/ # Explainable AI
│ ├── ai-agents/ # Building AI Agents
│ ├── computer-science/ # OOP & streaming pipelines
│ └── mlops-deployment/ # Planned module
│
├── 07-capstone-projects/ # 🎓 Projects
│ ├── industry-focus/ # Industry applications
│ ├── project-guides/ # Project guidelines
│ └── prerequisites/ # Prerequisites appendix
│
├── 08-data-science-applications-course/ # 🎯 Applied course
│ ├── presentation/ # Full course materials
│ ├── exercises/
│ └── assessments/ # Course assessments
│
├── 09-unit-tests/ # 🧪 Testing suites guide
│
├── assessments/ # 📝 Quizzes, exams, rubrics
├── datasets/ # 📦 Example datasets
├── docs/ # 📖 Guides and architecture notes
├── scripts/ # 🔧 Maintenance scripts
└── tests/ # ✅ Repository test suite
Per-module convention: each module directory contains presentation/ (Beamer slides), and where
applicable code/ (Python/R implementations) and exercises/ (problem sets).
1. Prerequisites — LaTeX, Python, R
LaTeX distribution:
# Ubuntu/Debian
sudo apt-get install texlive-full
# macOS
brew install --cask mactex
# Windows
# Download and install MiKTeX or TeX LivePython environment (for code examples):
pip install -r requirements.txt
# Or install the core set directly:
pip install numpy scipy matplotlib seaborn pandas scikit-learn statsmodels
pip install torch tensorflow # For deep learning examples
pip install shap lime # For XAI examplesA conda environment is also provided in environment.yml.
R environment (for R examples):
install.packages(c(
"AER", "rdrobust", "fixest", "did", # Causal inference
"caret", "recipes", "mice", # Feature engineering
"forecast", "vars", "fable" # Time series
))Or run the bundled installer: install_r_packages.R.
2. Compiling presentations — manual, latexmk, or CI
Manual compilation:
cd 02-deep-learning/deep-learning-fundamentals/presentation/
pdflatex deep_learning_beamer.tex
pdflatex deep_learning_beamer.tex # Run twice for referencesUsing latexmk (recommended):
cd 02-deep-learning/reinforcement-learning/presentation/
latexmk -pdf rl_beamer.texCompile everything:
./compile_all.shAutomated compilation:
- Push to GitHub → GitHub Actions automatically compiles all PDFs
- Download compiled PDFs from Actions artifacts or Releases
Build artifact policy:
- Compiled PDFs are tracked in git, so any deck can be read straight from GitHub without downloading a release or compiling it yourself.
- LaTeX auxiliary files (
.aux,.log,.fls,.fdb_latexmk,.nav,.snm,.out) are generated noise and are ignored. - CI additionally attaches freshly compiled PDFs to each release, so the release assets always reflect the latest source even if a tracked PDF is a commit behind.
3. Running code and exercises
Python:
# MCMC examples (if code/ directory exists with implementations)
# Example references are embedded in presentation materials
# Exercises and assessments
cd 03-bayesian-methods/mcmc/exercises/
pdflatex mcmc_exercises.texExercises:
# MCMC exercises
cd 03-bayesian-methods/mcmc/exercises/
pdflatex mcmc_exercises.tex
# Causal inference exercises
cd 04-causal-inference/causal-inference-fundamentals/exercises/
pdflatex causal_inference_exercises.texAll presentations use the ESMAD Beamer Theme for a consistent, professional appearance.
Theme features and usage template
✅ Professional color palette (ESMAD Blue, accents)
✅ Custom environments (theorems, definitions, examples, alerts)
✅ Mathematical notation helpers (\Normal, \E, \Var, etc.)
✅ Code listing styles with syntax highlighting
✅ Author information with ORCID integration
✅ Slide templates (title, TOC, contact, references)
✅ High-contrast variant for accessibility
\documentclass[aspectratio=169]{beamer}
\usepackage{../../../shared/theme/esmad_beamer_theme}
% Author info
\authorname{Your Name}
\authoremail{your.email@university.edu}
\authororcid{0000-0000-0000-0000}
\title{Your Presentation}
\date{\today}
\begin{document}
\begin{frame}
\titlepage
\end{frame}
% Your content...
\contactslide
\end{document}See shared/theme/STYLE_GUIDE.md for complete documentation, and
ACCESSIBILITY.md for accessibility guidance.
📖 For students — learning paths and study tips
Path 1: Machine Learning Fundamentals
- Statistical Learning (4 weeks)
- Feature Engineering (2 weeks)
- Optimization (3 weeks)
- Explainable AI (2 weeks)
Path 2: Deep Learning Specialization
- Deep Learning Fundamentals (4 weeks)
- Optimization (focus on neural networks)
- Reinforcement Learning (4 weeks)
- Time Series Analysis (focus on deep methods)
Path 3: Causal & Bayesian Methods
- Causal Inference (5 weeks)
- Bayesian ML (4 weeks)
- MCMC Methods (3 weeks)
- A/B Testing (2 weeks)
Competency matrices and additional paths live in docs/learning-paths/.
- 📚 Start with slides to understand concepts
- 💻 Run code examples to see methods in action
- 📝 Complete exercises to test understanding
- 📖 Read references for deeper knowledge
- 🤝 Join discussions (create GitHub issues)
👨🏫 For educators — course integration, customization, assessments
These materials can be integrated into:
- Graduate courses in Data Science/Statistics/CS
- Professional training programs
- Workshop series
- Seminar courses
- Fork this repository
- Customize presentations for your needs
- Add your own examples and exercises
- Maintain attribution (CC BY-SA 4.0)
Use the materials in assessments/:
- Quizzes for each topic
- Midterm and final exams
- Grading rubrics
- Project ideas
Per-topic enhancement guides are in docs/enhancement-guides/.
🔬 For researchers — citation and bibliographies
If you use these materials in your research or teaching, please cite:
@misc{ribeiro2025academic,
author = {Ribeiro, Diogo},
title = {Academic Presentations: Comprehensive Data Science Course Materials},
year = {2025},
publisher = {GitHub},
url = {https://github.com/diogoribeiro7/academic-presentations},
note = {Faculty of Media Arts and Design, Technical University of Porto}
}All presentations reference comprehensive BibTeX files:
\usepackage[backend=bibtex]{biblatex}
\addbibresource{../../../shared/bibliographies/mcmc_references.bib}
% In document
\cite{metropolis1953}
\cite{hoffman2014}
% At end
\printbibliographyAvailable:
shared/bibliographies/mcmc_references.bib: 30+ MCMC papersshared/bibliographies/causal_inference_references.bib: 50+ causal inference papersshared/bibliographies/statistical_learning_references.bib: 60+ ML/stats papers- Plus capstone, industry-focus, and per-topic enhancement bibliographies
All include DOIs for easy access.
CI/CD workflows
compile-latex.yml: Auto-compiles all LaTeX on pushcheck-links.yml: Verifies all URLs and DOIs weeklygenerate-previews.yml: Creates PDF preview gallerydependabot.yml: Keeps dependencies updated
Pre-commit hooks (formatting, spell check, LaTeX lint) are configured in
.pre-commit-config.yaml.
PDF preview gallery: https://diogoribeiro7.github.io/academic-presentations/
How to contribute
Contributions are welcome — see CONTRIBUTING.md for full guidelines.
- Fork the repository
- Create a feature branch
- Make your changes
- Test compilation and code
- Submit a pull request
Contribution types:
- 🐛 Fix errors in presentations
- 📚 Add new presentations
- 💡 Improve existing content
- 📖 Enhance documentation
- 🧪 Add code examples
- 📝 Create exercises
- 🎨 Improve theme/styling
Quality standards are documented in QUALITY.md.
License — CC BY-SA 4.0 for content, MIT for code
Licensed under Creative Commons Attribution-ShareAlike 4.0 International
You are free to:
- ✅ Share — copy and redistribute
- ✅ Adapt — remix, transform, and build upon
Under the terms:
- 📝 Attribution required
- 🔄 ShareAlike for derivatives
Code examples licensed under MIT License
Contact & collaboration
- Email: dfr@esmad.ipp.pt
- Institution: Faculty of Media Arts and Design, Technical University of Porto
- LinkedIn: diogo-ribeiro-9094604a
- ORCID: 0009-0001-2022-7072
- Markov Chain Monte Carlo and Bayesian computation
- Machine learning and deep learning
- Causal inference and econometrics
- Financial risk modeling
- Time series analysis and forecasting
- 🎓 Guest lectures and workshops
- 🏢 Corporate training programs
- 🔬 Research collaborations
- 📝 Joint publications
- 🌐 Conference presentations
Acknowledgments
- Students and colleagues for valuable feedback
- Open source community for tools and inspiration
- Academic community for rigorous peer review
Repository maintainer: Diogo Ribeiro · Status: ✅ Actively maintained · History: CHANGELOG.md · View releases