Six business-flavoured A/B-testing and experiment-design cases (with slide decks) plus a 12-part statistical-inference workbook — from descriptive statistics to power analysis and MDE, all in runnable notebooks.
The lead exhibit — the same dataset "proving" two opposite conclusions:
What you're looking at: both panels are built from the same 500 users. Pick 200 of them one way — product A "wins" 216 vs 30; pick another 200 — product B "wins" 221 vs 29. Nothing was faked inside either panel; only the choice of who gets counted changed. The notebook then shows the honest method — proportional stratified sampling — that makes the answer stop depending on the picker.
Statistics is the discipline of not fooling yourself with data — and these notebooks show both sides: how numbers lie when collected carelessly, and how to design measurements so they can't.
- A p-value is "how surprised an honest skeptic should be": if the new ad actually changed nothing, how often would a difference this big appear by pure luck? Small p — the skeptic is surprised — maybe something real. p = 0.995 (one of the cases) — the data is utterly unsurprising, and the honest verdict is "don't ship".
- Power and MDE answer the question before the experiment: how many users do we need so that, if the effect is real, we actually catch it — and what's the smallest effect worth catching?
- Bonferroni correction: run 21 comparisons and one will look "significant" by luck alone — like buying 21 lottery tickets and being amazed one won. The bar must rise with the number of tests.
Every notebook has a companion beginner's guide (.md file next to it) explaining the case from zero, with analogies and a glossary.
| Case | Question | Punchline |
|---|---|---|
| Ad campaign ROI · guide · deck | Which of two ad campaigns to scale? | +884 incremental purchases → ₽2.65M revenue vs ₽1.54M cost; target the 23+ segments |
| Selection bias · guide · deck | Can sampling "prove" anything? | Yes — both directions, from the same data; fixed with stratified sampling |
| Customer targeting & Bonferroni · guide · deck | Which segment is worth acquiring? | Mean profit ₽16,515 vs ₽15,000 CAC; n=248 needed for significance; 21 pairwise age tests → α_adj≈0.0024, so aggregate age buckets; recommendation: women 18–24 |
| A/B design & MDE · guide · deck | Design a test end-to-end | Required n=1,916/group, only 1,356 available → underpowered; observed p=0.995 → honest "do not ship" |
| Rollout false-positive risk · guide · deck | Is "200 visits, ship if conv ≥21%" safe? | No: 15.6% false-positive rollout risk; 1,000 impressions cut it to 3.7% |
| Hiring-test threshold · guide · deck | Where to set a screening-test bar? | 19/20 correct: strong candidates pass ≈85%, weak <50%; plus test-length sweep |
| # | Notebook | Topic |
|---|---|---|
| 01 | Descriptive statistics · guide | Pivot tables, data cleaning, segment profitability |
| 02 | Sampling & populations · guide | Population vs sample, bias scenarios |
| 03 | Joint distributions · guide | 2-D discrete distributions, independence |
| 04 | Sampling distributions · guide | Behaviour of the mean, stratification |
| 05 | Discrete random variables · guide | CDFs, transforms, sums — nine worked problems |
| 06 | Binomial & continuous models · guide | Pass-probability curves, E/V derivations |
| 07 | Normal distribution · guide | z-scores, normal approximation |
| 08 | CLT & sample size · guide | Insurance-risk portfolio, minimum n |
| 09 | Confidence intervals · guide | Monte-Carlo coverage, estimator properties |
| 10 | Hypothesis formulation · guide | Valid vs invalid H₀/H₁ |
| 11 | Binomial tests, power & MDE · guide | Screening test end-to-end: α, power, MDE |
| 12 | Two-sample tests · guide | z/t-tests, unequal variances, power |
Above: from workbook notebook 09 — hundreds of simulated confidence intervals; the few that miss the true value (they should be ~1% at the 99% level) are exactly what "confidence level" means.
pip install -r requirements.txt
jupyter labDatasets auto-download from a public GitHub source; everything runs on CPU. Decks in docs/decks/ are in Russian.
Keywords: statistics, A/B testing, hypothesis testing, statistical power, MDE, Bonferroni correction, confidence intervals, experiment design, Monte Carlo, business analytics
Ключевые слова: статистика, A/B тестирование, проверка гипотез, мощность критерия, MDE, поправка Бонферрони, доверительные интервалы, дизайн экспериментов, бизнес-аналитика

