Skip to content

Latest commit

 

History

2 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Applied Statistics Casebook

Six business-flavoured A/B-testing and experiment-design cases (with slide decks) plus a 12-part statistical-inference workbook — from descriptive statistics to power analysis and MDE, all in runnable notebooks.

Русская версия

Python SciPy statsmodels License

The lead exhibit — the same dataset "proving" two opposite conclusions:

selection bias demo

What you're looking at: both panels are built from the same 500 users. Pick 200 of them one way — product A "wins" 216 vs 30; pick another 200 — product B "wins" 221 vs 29. Nothing was faked inside either panel; only the choice of who gets counted changed. The notebook then shows the honest method — proportional stratified sampling — that makes the answer stop depending on the picker.

In plain words — from zero

Statistics is the discipline of not fooling yourself with data — and these notebooks show both sides: how numbers lie when collected carelessly, and how to design measurements so they can't.

  • A p-value is "how surprised an honest skeptic should be": if the new ad actually changed nothing, how often would a difference this big appear by pure luck? Small p — the skeptic is surprised — maybe something real. p = 0.995 (one of the cases) — the data is utterly unsurprising, and the honest verdict is "don't ship".
  • Power and MDE answer the question before the experiment: how many users do we need so that, if the effect is real, we actually catch it — and what's the smallest effect worth catching?
  • Bonferroni correction: run 21 comparisons and one will look "significant" by luck alone — like buying 21 lottery tickets and being amazed one won. The bar must rise with the number of tests.

Every notebook has a companion beginner's guide (.md file next to it) explaining the case from zero, with analogies and a glossary.

A/B testing & experiment design cases — ab-testing-cases/

Case Question Punchline
Ad campaign ROI · guide · deck Which of two ad campaigns to scale? +884 incremental purchases → ₽2.65M revenue vs ₽1.54M cost; target the 23+ segments
Selection bias · guide · deck Can sampling "prove" anything? Yes — both directions, from the same data; fixed with stratified sampling
Customer targeting & Bonferroni · guide · deck Which segment is worth acquiring? Mean profit ₽16,515 vs ₽15,000 CAC; n=248 needed for significance; 21 pairwise age tests → α_adj≈0.0024, so aggregate age buckets; recommendation: women 18–24
A/B design & MDE · guide · deck Design a test end-to-end Required n=1,916/group, only 1,356 available → underpowered; observed p=0.995 → honest "do not ship"
Rollout false-positive risk · guide · deck Is "200 visits, ship if conv ≥21%" safe? No: 15.6% false-positive rollout risk; 1,000 impressions cut it to 3.7%
Hiring-test threshold · guide · deck Where to set a screening-test bar? 19/20 correct: strong candidates pass ≈85%, weak <50%; plus test-length sweep

Statistical inference workbook — inference-workbook/

# Notebook Topic
01 Descriptive statistics · guide Pivot tables, data cleaning, segment profitability
02 Sampling & populations · guide Population vs sample, bias scenarios
03 Joint distributions · guide 2-D discrete distributions, independence
04 Sampling distributions · guide Behaviour of the mean, stratification
05 Discrete random variables · guide CDFs, transforms, sums — nine worked problems
06 Binomial & continuous models · guide Pass-probability curves, E/V derivations
07 Normal distribution · guide z-scores, normal approximation
08 CLT & sample size · guide Insurance-risk portfolio, minimum n
09 Confidence intervals · guide Monte-Carlo coverage, estimator properties
10 Hypothesis formulation · guide Valid vs invalid H₀/H₁
11 Binomial tests, power & MDE · guide Screening test end-to-end: α, power, MDE
12 Two-sample tests · guide z/t-tests, unequal variances, power

coverage

Above: from workbook notebook 09 — hundreds of simulated confidence intervals; the few that miss the true value (they should be ~1% at the 99% level) are exactly what "confidence level" means.

Getting started

pip install -r requirements.txt
jupyter lab

Datasets auto-download from a public GitHub source; everything runs on CPU. Decks in docs/decks/ are in Russian.


Keywords: statistics, A/B testing, hypothesis testing, statistical power, MDE, Bonferroni correction, confidence intervals, experiment design, Monte Carlo, business analytics

Ключевые слова: статистика, A/B тестирование, проверка гипотез, мощность критерия, MDE, поправка Бонферрони, доверительные интервалы, дизайн экспериментов, бизнес-аналитика

About

Applied statistics casebook: A/B-testing business cases (ROI, MDE, Bonferroni, selection bias) with decks, plus a 12-part statistical inference workbook

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages