Articles on statistical power

A Comprehensive Guide to Observed Power (Post Hoc Power)

“Observed power”, “Post hoc Power”, and “Retrospective power” all refer to the statistical power of a statistical significance test to detect a true effect equal to the observed effect. In a broader sense these terms may also describe any power analysis performed after an experiment has completed. Importantly, it is the first, narrower sense that […] Read more…

Posted in Statistics | Also tagged mde, minimum detectable effect, minimum effect of interest, observed power, post hoc power

What if the Observed Effect is Smaller Than the MDE?

The above is a question asked by some practitioners of A/B testing, as well as a number of their clients when examining the outcome of an online controlled experiment. It may be raised regardless if the outcome is statistically significant or not. In both cases the fact the observed effect in an A/B test is […] Read more…

Posted in A/B testing, Statistics | Also tagged mde, minimum detectable effect, minimum effect of interest, observed power

Using Observed Power in Online A/B Tests

Observed power, often referred to as “post hoc power” and “retrospective power” is the statistical power of a test to detect a true effect equal to the observed effect size. “Detect” in the context of a statistical hypothesis test means to result in a statistically significant outcome. Some calculators aimed at A/B testing practitioners use […] Read more…

Posted in A/B testing, Statistics | Also tagged observed power, optional stopping, peeking, post hoc power, underpowred tests

Stop AbUsing the Mann-Whitney U Test (MWU)

The Mann Whitney U Test (MWU), also known as the Wilcoxon Rank Sum Test and the Mann-Whitney-Wilcoxon Test, continues to be advertised as the go-to test for analyzing non-normally distributed data. In online experimentation it is often touted as the most suitable for analyses of non-binomial metrics with typically non-normal (skewed) distributions such as average […] Read more…

Posted in A/B testing, Statistics | Also tagged arpu, average revenue per user, difference in medians, mann-whitney u test, mwu, skewed distribution, skewness, statistical significance, stochastic difference

False Positive Risk in A/B Testing

Have you heard how there is a much greater probability than generally expected that a statistically significant test outcome is in fact a false positive? In industry jargon: that a variant has been identified as a “winner” when it is not. In demonstrating the above the terms “False Positive Risk” (FPR), “False Findings Rate” (FFR), […] Read more…

Posted in A/B testing, Bayesian A/B testing, Statistics | Also tagged bayes rule, bayesian inference, false discovery rate, false findings rate, false positive rate, false positive risk, fpr, p-value, type I error

How to Run Shorter A/B Tests?

Running shorter tests is key to improving the efficiency of experimentation as it translates to smaller direct losses from testing inferior experiences and also less unrealized revenue due to late implementation of superior ones. Despite this, many practitioners are yet to start conducting tests at the frontier of efficiency. This article presents ways to shorten […] Read more…

Posted in A/B testing, Statistics | Also tagged ab testing, efficient testing, small sample size, test duration

Comparison of the statistical power of sequential tests: SPRT, AGILE, and Always Valid Inference

Power and Average Sample Size of Sequential Tests

In A/B testing sequential tests are gradually becoming the norm due to the increased efficiency and flexibility that they grant practitioners. In most practical scenarios sequential tests offer a balance of risks and rewards superior to that of an equivalent fixed sample test. Sequential monitoring achieves this superiority by trading statistical power for the ability […] Read more…

Posted in A/B testing, AGILE A/B testing, Statistics | Also tagged msprt, power and sample size, sample size, sequential testing, sequential tests, sprt

Statistical Power, MDE, and Designing Statistical Tests

One topic has surfaced in my ten years of developing statistical tools, consulting, and participating in discussions and conversations with CRO & A/B testing practitioners as causing the most confusion and that is statistical power and the related concept of minimum detectable effect (MDE). Some myths were previously dispelled in “Underpowered A/B tests – confusions, […] Read more…

Posted in A/B testing, Statistics | Also tagged minimum detectable effect, minimum effect of interest, risk reward analysis, risk-reward ratio

What Can Be Learned From 1,001 A/B Tests?

How long does a typical A/B test run for? What percentage of A/B tests result in a ‘winner’? What is the average lift achieved in online controlled experiments? How good are top conversion rate optimization specialists at coming up with impactful interventions for websites and mobile apps? This meta-analysis of 1,001 A/B tests analyzed using […] Read more…

Posted in A/B testing, AGILE A/B testing, Conversion optimization | Also tagged conversion rate optimization, efficient ab testing, lift, meta analysis, minimum detectable effect, minimum effect of interest, online ab testing, sequential testing, test duration

Fully Sequential vs Group Sequential Tests

What is the best design for a statistical test with sequential evaluation of the data at multiple points in time? This is a question anyone who has realized that unaccounted for peeking with intent to stop is the bane of A/B testing eventually comes to ask. So how does one go about answering that? This […] Read more…

Posted in AGILE A/B testing, Statistics | Also tagged ab testing, conversion rate optimization, sequential testing, sequential tests, statistical significance

A/B Testing Statistics – A Concise Guide for Non-Statisticians

Navigating the maze of A/B testing statistics can be challenging. This is especially true for those new to statistics and probability. One reason is the obscure terminology popping up in every other sentence. Another is that the writings can be vague, conflicting, incomplete, or simply wrong, depending on the source. Articles sprinkled with advanced math, […] Read more…

Posted in A/B testing, Statistics | Also tagged ab testing, confidence intervals, p-value, statistical confidence, statistical significance

Underpowered A/B Tests – Confusions, Myths, and Reality

In recent years a lot more CRO & A/B testing practitioners have started paying more attention to the statistical power of their online experiments, at least based on my observations. While this a positive development for which I hope I had contributed somewhat, it comes with the inevitable confusions and misunderstandings surrounding a complex concept […] Read more…

Posted in A/B testing, Statistics | Also tagged ab testing methodology, minimum detectable effect, minimum effect of interest, statistical test design, type II error, underpowered

Search

Browse by topic

Browse by year

The book on user testing

Take your A/B testing program to the next level

Learn more

Articles on statistical power

Search

Recent articles

Browse by topic

Browse by year