Adaptive Bayesian Clinical Trial
-
Updated
Jul 20, 2020 - R
Adaptive Bayesian Clinical Trial
Statistical power analyses in the browser
Power and Sample Size Calculation for the Cochran-Mantel-Haenszel Chi-Squared Test
Code for "Adaptive Selection of the Optimal Strategy to Improve Precision and Power in Randomized Trials"
PRISME Power Calculator
Find out which qualities of your writing actually predict engagement. Rates every post you have published against a pre-registered rubric using Jev's calibrated judgments, then tests those ratings against your real engagement numbers. Refuses to report findings your sample cannot support.
This incomplete repository is used to facilitate the consultation of individual files in this project. Only files smaller than 100 MB are available here. The complete project is available at https://doi.org/10.17605/OSF.IO/GT5UF.
Assay-aware observability, donor-level power and design adequacy for single-cell alternative splicing
How many runs before your eval means anything? Reliability statistics for stochastic evals: audit miss rates, exact intervals, runs-needed.
A probe suite that measures which conversation states an LLM cannot leave. Three arms, because two cannot tell obedience from token statistics; a null only counts when the design had the power to see the effect.
A/B test analyzer that returns ship/hold/iterate/kill, not a p-value. Power vs a pre-specified MDE, CIs, effect size, and SRM checks on every result. 37 tests.
Eval suites that tell you when they've gone blind: coverage, detection power and judge depth for LLM agent evaluation.
Identifying and avoiding common misinterpretations in using statistics
Applied statistics casebook: A/B-testing business cases (ROI, MDE, Bonferroni, selection bias) with decks, plus a 12-part statistical inference workbook
Simulation studies of power and Type I error of mass univariate statistics for ERP data
Size your early-stopping window by statistical power instead of by habit
Your prompt eval cannot detect what you think it can. Ship-the-higher-number declares a winner 46.6% of the time on identical variants; detecting +5pp at 80% power needs ~859 items. Measured by simulation, re-measured in CI.
What a 1,450-test tail-predictor search could have detected: a power accounting, and a tail-IC/mean-IC pre-registration gate. Companion to Yan (2026).
Audits a data quality suite for rules that are statistically incapable of catching the defect they were written to catch, and for the damage a block-sampling clause does to the ones that could. 15 of 40 rules cannot fail; 13 of those block the pipeline; on a block-sampling warehouse it rises to 28.
To associate your repository with the statistical-power topic, visit your repo's landing page and select "manage topics."