A supervised learning analysis of the IBM Telco Customer Churn dataset, applying classification and regression methods to predict customer churn and cumulative revenue. Developed as the final project for the Advanced Modelling course, Master in Computational Social Science, Universidad Carlos III de Madrid (2025–2026).
├── Final_Project.qmd # Main Quarto document (source)
├── Final_Project.html # Rendered HTML report (self-contained)
├── Telco_Customer_Churn.csv # Dataset
└── README.md # This file
This project applies a full supervised learning pipeline to two parallel analytical tasks:
Classification — Predicting which customers will churn, with emphasis on class imbalance handling and cost-sensitive learning. Methods include logistic regression with stepwise selection, LDA, QDA, k-Nearest Neighbours, Decision Trees, and Random Forests. A custom cost matrix (FN=50, FP=10) is used to derive a business-optimal classification threshold.
Regression — Predicting log(TotalCharges) as a proxy for customer lifetime value, using service subscriptions, contract type, and demographics as predictors. Methods include OLS, Ridge, Lasso, Elastic Net, PCR, PLS, Regression Trees, and Random Forests.
- Contract type is the single strongest predictor in both tasks — month-to-month customers are the highest churn risk and generate the least cumulative revenue
- Fiber optic internet customers churn at 7.5× the rate of DSL customers
- Threshold adjustment (0.5 → 0.30) reduces expected misclassification cost by 31% before any model change
- Resampling (downsampling/upsampling) achieves the lowest expected cost when retention capacity is unconstrained
- Regularisation (Ridge, Lasso, Elastic Net) offers no improvement over OLS — the feature set is already parsimonious
- Random Forest is the only method to meaningfully beat OLS in regression (RMSE 0.847 vs 0.888)
The analysis is written in R and rendered with Quarto. The following packages are required:
install.packages(c(
"tidyverse", "dplyr", "caret", "MASS", "e1071", "class",
"tree", "randomForest", "glmnet", "pls", "pROC",
"GGally", "patchwork"
))IBM Telco Customer Churn Dataset - 7,043 customer records, 21
variables - Binary classification target: Churn (Yes/No) - Regression
target: log(TotalCharges)
Sources: - Kaggle - IBM Community
- Clone the repository
- Place
Telco_Customer_Churn.csvin the root directory - Open
Final_Project.qmdin RStudio - Click Render or run:
quarto::quarto_render("Final_Project.qmd")The output Final_Project.html is fully self-contained and can be
opened in any browser without additional dependencies.
Reichheld, F. F., & Schefter, P. (2000). E-loyalty: Your secret weapon on the Web. Harvard Business Review, 78(4), 105–113.
IBM. (2019). Telco Customer Churn Dataset. IBM Community.
Tommaso Accornero Master in Computational Social Science — Universidad Carlos III de Madrid