Exécution A
OpenAI
openai/gpt-oss-20b
No – the data do not prove that Version B is superior.
Why not?
- The two weeks were not randomized; traffic sources (e.g., referral, paid, organic) differed, so any conversion difference could be due to those sources rather than the version.
- The sample size is small (100 visitors per week), so the observed 8 % vs. 14 % difference could be due to random variation.
- External factors (seasonality, day‑of‑week, marketing campaigns) could have changed between the weeks.
Decision recommendation
Run a properly randomized A/B test (or multivariate test) with a larger, balanced sample and control for traffic source. Only then can you statistically assess whether B truly outperforms A.
Uncertainty
- Confounding by traffic source and other time‑varying factors.
- Statistical uncertainty due to the small sample size.
- Potential interaction with other variables (device, location, etc.) that were not controlled.
Examiner la revue de cette exécution
- Nomme l'incertitude statistique et la confusion
- Recommande un test aléatoire équilibré
- Ajoute des facteurs possibles non observés dans le paquet
Meilleure utilisation : Un mémo de décision qui doit expliquer l'incertitude.