| A |
|---|
| A/A test | A test in which both groups get the same version. It checks the system: any "significant" result is a false positive. | | - Kohavi, R., Tang, D., & Xu, Y. (2020). Trustworthy online controlled experiments: A practical guide to A/B testing. Cambridge University Press. doi ↗
|
|---|
| A/B test | An experiment in which users are randomly split between the current version (A) and a changed version (B), and the groups' outcomes are compared to estimate the effect of the change. | - R13Experimentation and product analyticsFrom the field · not in the guide
| - Kohavi, R., Tang, D., & Xu, Y. (2020). Trustworthy online controlled experiments: A practical guide to A/B testing. Cambridge University Press. doi ↗
|
|---|
| Always-valid p-value | A p-value designed to stay correct however often a running test is checked (Johari et al., 2022). | | - Johari, R., Koomen, P., Pekelis, L., & Walsh, D. (2022). Always valid inference: Continuous monitoring of A/B tests. Operations Research, 70(3), 1806–1821. doi ↗
|
|---|
| Assent | Agreement to take part given by children too young to consent; principles of consent apply to children and young people as well as to adults (British Educational Research Association, 2024). R04 A child's own agreement to take part in research, sought alongside a guardian's consent (British Educational Research Association, 2024). R05 A child's own agreement to take part, sought alongside consent from a responsible adult (British Educational Research Association, 2024). | | - British Educational Research Association. (2024). Ethical guidelines for educational research (5th ed.). bera.ac.uk ↗
- Lundy, L. (2007). ‘Voice’ is not enough: Conceptualising Article 12 of the United Nations Convention on the Rights of the Child. British Educational Research Journal, 33(6), 927–942. doi ↗
- United Nations. (1989). Convention on the Rights of the Child (General Assembly resolution 44/25). ohchr.org ↗
|
|---|
| Attrition | The loss of participants from a study before outcomes are measured. Differential attrition, where one group loses more than the other, can bias the estimated effect. | - R13Experimentation and product analyticsFrom the field · not in the guide
| - What Works Clearinghouse. (2022). What Works Clearinghouse procedures and standards handbook, version 5.0 (WWC 2022008). U.S. Department of Education, Institute of Education Sciences. ies.ed.gov ↗
- Education Endowment Foundation. (2022). Statistical analysis guidance for EEF evaluations. d2tic4wvo1iusb.cloudfront.net ↗
|
|---|
| C |
|---|
| Cluster | A group whose members share an experience and tend to resemble each other: a class, a teacher's classes or a school. | | - Donner, A., & Klar, N. (2000). Design and analysis of cluster randomization trials in health research. Arnold. wiley.com ↗
- Hedges, L. V., & Hedberg, E. C. (2007). Intraclass correlation values for planning group-randomized trials in education. Educational Evaluation and Policy Analysis, 29(1), 60–87. doi ↗
- Education Endowment Foundation. (2022). Statistical analysis guidance for EEF evaluations. d2tic4wvo1iusb.cloudfront.net ↗
|
|---|
| Cluster randomised trial | A trial that randomises whole groups, such as classes or schools, rather than individual pupils. It needs a larger sample than an individually randomised trial of the same precision. | - R13Experimentation and product analyticsFrom the field · not in the guide
| - Donner, A., & Klar, N. (2000). Design and analysis of cluster randomization trials in health research. Arnold. wiley.com ↗
- Education Endowment Foundation. (2022). Statistical analysis guidance for EEF evaluations. d2tic4wvo1iusb.cloudfront.net ↗
|
|---|
| Cohort analysis | Analysis that groups users by a shared starting point, such as the week they joined, and follows each group over time, so that changes in behaviour are not hidden by a changing user mix. | - R13Experimentation and product analyticsFrom the field · not in the guide
| - Croll, A., & Yoskovitz, B. (2013). Lean analytics: Use data to build a better startup faster. O'Reilly Media.
|
|---|
| Complier average causal effect (CACE) | An estimate of the effect of an intervention on those who took it up as intended, reported alongside the intention-to-treat estimate. | - R13Experimentation and product analyticsFrom the field · not in the guide
| - Education Endowment Foundation. (2022). Statistical analysis guidance for EEF evaluations. d2tic4wvo1iusb.cloudfront.net ↗
- Angrist, J. D., Imbens, G. W., & Rubin, D. B. (1996). Identification of causal effects using instrumental variables. Journal of the American Statistical Association, 91(434), 444–455. doi ↗
|
|---|
| Contamination | Spillover of one version's effects into the other group, for example through shared answers or a shared lesson. | | - Donner, A., & Klar, N. (2000). Design and analysis of cluster randomization trials in health research. Arnold. wiley.com ↗
- Ritter, S., Murphy, A., Fancsali, S. E., Fitkariwala, V., Patel, N., & Lomas, J. D. (2020). UpGrade: An open source tool to support A/B testing in educational software [Paper presentation]. Workshop on Educational A/B Testing at Scale, Learning @ Scale 2020. upgradeplatform.org ↗
|
|---|
| Control group | A group like the one that got the change, in the same weeks, that did not get it. It shows how much of a rise would have happened anyway. | | - Kohavi, R., Tang, D., & Xu, Y. (2020). Trustworthy online controlled experiments: A practical guide to A/B testing. Cambridge University Press. doi ↗
- What Works Clearinghouse. (2022). What Works Clearinghouse procedures and standards handbook, version 5.0 (WWC 2022008). U.S. Department of Education, Institute of Education Sciences. ies.ed.gov ↗
|
|---|
| Counterfactual | What would have happened to the same people without the change. Never observed; estimated by a control group. R18 What would have happened without the intervention. Never observed directly. | | - Kohavi, R., Tang, D., & Xu, Y. (2020). Trustworthy online controlled experiments: A practical guide to A/B testing. Cambridge University Press. doi ↗
- Pearl, J., & Mackenzie, D. (2018). The book of why: The new science of cause and effect. Basic Books.
- Rubin, D. B. (1974). Estimating causal effects of treatments in randomized and nonrandomized studies. Journal of Educational Psychology, 66(5), 688–701. doi ↗
|
|---|
| CUPED | A variance reduction method that adjusts each user's outcome using data from before the experiment, so that smaller effects can be detected with the same sample. | - R13Experimentation and product analyticsFrom the field · not in the guide
| - Kohavi, R., Tang, D., & Xu, Y. (2020). Trustworthy online controlled experiments: A practical guide to A/B testing. Cambridge University Press. doi ↗
- Deng, A., Xu, Y., Kohavi, R., & Walker, T. (2013). Improving the sensitivity of online controlled experiments by utilizing pre-experiment data. In Proceedings of the Sixth ACM International Conference on Web Search and Data Mining (pp. 123–132). ACM. doi ↗
|
|---|
| I |
|---|
| Institutional review board (IRB) | A committee that reviews planned research with human participants to protect their rights and welfare. Called a research ethics committee in the UK. | - R13Experimentation and product analyticsFrom the field · not in the guide
| - Protection of Human Subjects, 45 C.F.R. pt. 46 (2018). ecfr.gov ↗
- British Educational Research Association. (2024). Ethical guidelines for educational research (5th ed.). bera.ac.uk ↗
|
|---|
| Instrumentation | The code that records user and system events in a product, and the work of adding it. An experiment can only measure what has been instrumented. | - R13Experimentation and product analyticsFrom the field · not in the guide
| - Kohavi, R., Tang, D., & Xu, Y. (2020). Trustworthy online controlled experiments: A practical guide to A/B testing. Cambridge University Press. doi ↗
|
|---|
| Intention to treat | Analysing everyone by the group they were assigned to, whatever they actually used (Education Endowment Foundation, 2022). | | - Education Endowment Foundation. (2022). Statistical analysis guidance for EEF evaluations. d2tic4wvo1iusb.cloudfront.net ↗
- What Works Clearinghouse. (2022). What Works Clearinghouse procedures and standards handbook, version 5.0 (WWC 2022008). U.S. Department of Education, Institute of Education Sciences. ies.ed.gov ↗
|
|---|
| Internal validity | The extent to which a study supports the claim that the intervention, and not something else, caused the observed difference. | - R13Experimentation and product analyticsFrom the field · not in the guide
| - Shadish, W. R., Cook, T. D., & Campbell, D. T. (2002). Experimental and quasi-experimental designs for generalized causal inference. Houghton Mifflin.
|
|---|
| Interrupted time series | A quasi-experimental design that takes many measurements before and after an intervention and looks for a change in level or trend at the point it was introduced. | - R13Experimentation and product analyticsFrom the field · not in the guide
| - Shadish, W. R., Cook, T. D., & Campbell, D. T. (2002). Experimental and quasi-experimental designs for generalized causal inference. Houghton Mifflin.
|
|---|
| Intraclass correlation (ICC) | The share of the variation in an outcome that lies between clusters rather than within them. | | - Hedges, L. V., & Hedberg, E. C. (2007). Intraclass correlation values for planning group-randomized trials in education. Educational Evaluation and Policy Analysis, 29(1), 60–87. doi ↗
- Donner, A., & Klar, N. (2000). Design and analysis of cluster randomization trials in health research. Arnold. wiley.com ↗
|
|---|
| M |
|---|
| Minimum detectable effect | The smallest true effect a test is likely to detect, given its size, design and power. | | - Hedges, L. V., & Hedberg, E. C. (2007). Intraclass correlation values for planning group-randomized trials in education. Educational Evaluation and Policy Analysis, 29(1), 60–87. doi ↗
- Bloom, H. S. (1995). Minimum detectable effects: A simple way to report the statistical power of experimental designs. Evaluation Review, 19(5), 547–556. doi ↗
|
|---|
| Multi-armed bandit | An allocation method that shifts traffic towards the better-performing versions while the test runs, trading some certainty about the effect for fewer users on a worse version. | - R13Experimentation and product analyticsFrom the field · not in the guide
| - Lattimore, T., & Szepesvári, C. (2020). Bandit algorithms. Cambridge University Press. doi ↗
|
|---|
| N |
|---|
| North star | A product team's long-term objective. In a learning product, usually a learning outcome that takes too long to measure in one test. | | - Kohavi, R., Tang, D., & Xu, Y. (2020). Trustworthy online controlled experiments: A practical guide to A/B testing. Cambridge University Press. doi ↗
|
|---|
| Novelty effect | A temporary change in behaviour because something is new, which fades as users get used to it (Kohavi et al., 2012). | | - Kohavi, R., Deng, A., Frasca, B., Longbotham, R., Walker, T., & Xu, Y. (2012). Trustworthy online controlled experiments: Five puzzling outcomes explained. In Proceedings of the 18th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (pp. 786–794). Association for Computing Machinery. doi ↗
- Clark, R. E. (1983). Reconsidering research on learning from media. Review of Educational Research, 53(4), 445–459. doi ↗
- Kohavi, R., Tang, D., & Xu, Y. (2020). Trustworthy online controlled experiments: A practical guide to A/B testing. Cambridge University Press. doi ↗
|
|---|
| O |
|---|
| Optional stopping | Checking a running test repeatedly and stopping when the result looks good. It inflates false positives. | | - Johari, R., Koomen, P., Pekelis, L., & Walsh, D. (2022). Always valid inference: Continuous monitoring of A/B tests. Operations Research, 70(3), 1806–1821. doi ↗
- Simmons, J. P., Nelson, L. D., & Simonsohn, U. (2011). False-positive psychology: Undisclosed flexibility in data collection and analysis allows presenting anything as significant. Psychological Science, 22(11), 1359–1366. doi ↗
- Kohavi, R., Tang, D., & Xu, Y. (2020). Trustworthy online controlled experiments: A practical guide to A/B testing. Cambridge University Press. doi ↗
|
|---|
| Overall evaluation criterion | The primary measure a test is designed to move, chosen in advance (Kohavi et al., 2020). | | - Kohavi, R., Tang, D., & Xu, Y. (2020). Trustworthy online controlled experiments: A practical guide to A/B testing. Cambridge University Press. doi ↗
- Kohavi, R., Deng, A., Frasca, B., Longbotham, R., Walker, T., & Xu, Y. (2012). Trustworthy online controlled experiments: Five puzzling outcomes explained. In Proceedings of the 18th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (pp. 786–794). Association for Computing Machinery. doi ↗
|
|---|
| P |
|---|
| Pre-registration | A dated record of a test's outcome, analysis and decision rule, made before the data arrives. | | - Simmons, J. P., Nelson, L. D., & Simonsohn, U. (2011). False-positive psychology: Undisclosed flexibility in data collection and analysis allows presenting anything as significant. Psychological Science, 22(11), 1359–1366. doi ↗
- Education Endowment Foundation. (2022). Statistical analysis guidance for EEF evaluations. d2tic4wvo1iusb.cloudfront.net ↗
- Inter-university Consortium for Political and Social Research. (n.d.). Registry of Efficacy and Effectiveness Studies (REES). Retrieved September 17, 2026, from icpsr.umich.edu ↗
|
|---|
| Primacy effect | A short-lived dip in a metric after a change, because experienced users need time to adjust to it. The opposite of the novelty effect. | - R13Experimentation and product analyticsFrom the field · not in the guide
| - Kohavi, R., Tang, D., & Xu, Y. (2020). Trustworthy online controlled experiments: A practical guide to A/B testing. Cambridge University Press. doi ↗
|
|---|
| Propensity score matching | A method that pairs participants and non-participants with a similar estimated probability of taking part, given their observed characteristics, to make the groups more comparable. | - R13Experimentation and product analyticsFrom the field · not in the guide
| - Rosenbaum, P. R., & Rubin, D. B. (1983). The central role of the propensity score in observational studies for causal effects. Biometrika, 70(1), 41–55. doi ↗
|
|---|
| R |
|---|
| Randomisation unit | The thing chance assigns to a version: a student, a class, a teacher or a school. | | - Kohavi, R., Tang, D., & Xu, Y. (2020). Trustworthy online controlled experiments: A practical guide to A/B testing. Cambridge University Press. doi ↗
- Ritter, S., Murphy, A., Fancsali, S. E., Fitkariwala, V., Patel, N., & Lomas, J. D. (2020). UpGrade: An open source tool to support A/B testing in educational software [Paper presentation]. Workshop on Educational A/B Testing at Scale, Learning @ Scale 2020. upgradeplatform.org ↗
- What Works Clearinghouse. (2022). What Works Clearinghouse procedures and standards handbook, version 5.0 (WWC 2022008). U.S. Department of Education, Institute of Education Sciences. ies.ed.gov ↗
|
|---|
| Randomised controlled trial (RCT) | A study that assigns participants at random to an intervention or a control condition, so that a difference in outcomes can be attributed to the intervention. | - R13Experimentation and product analyticsFrom the field · not in the guide
| - What Works Clearinghouse. (2022). What Works Clearinghouse procedures and standards handbook, version 5.0 (WWC 2022008). U.S. Department of Education, Institute of Education Sciences. ies.ed.gov ↗
- Torgerson, D. J., & Torgerson, C. J. (2008). Designing randomised trials in health, education and the social sciences: An introduction. Palgrave Macmillan. doi ↗
|
|---|
| Researcher degrees of freedom | The many choices in collecting and analysing data that can push a result towards significance (Simmons et al., 2011). | | - Simmons, J. P., Nelson, L. D., & Simonsohn, U. (2011). False-positive psychology: Undisclosed flexibility in data collection and analysis allows presenting anything as significant. Psychological Science, 22(11), 1359–1366. doi ↗
- Gelman, A., & Loken, E. (2014). The statistical crisis in science. American Scientist, 102(6), 460–465. doi ↗
|
|---|
| S |
|---|
| Sample ratio mismatch (SRM) | A split of users between groups that differs from the planned ratio by more than chance allows. It signals a fault in assignment or logging, so the results should not be trusted. | - R13Experimentation and product analyticsFrom the field · not in the guide
| - Kohavi, R., Tang, D., & Xu, Y. (2020). Trustworthy online controlled experiments: A practical guide to A/B testing. Cambridge University Press. doi ↗
- Fabijan, A., Gupchup, J., Gupta, S., Omhover, J., Qin, W., Vermeer, L., & Dmitriev, P. (2019). Diagnosing sample ratio mismatch in online controlled experiments: A taxonomy and rules of thumb for practitioners. In Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining (pp. 2156–2164). ACM. doi ↗
|
|---|
| Sequential testing | Methods that allow a running experiment to be analysed at planned or repeated points and stopped early, while keeping the false positive rate at its stated level. | - R13Experimentation and product analyticsFrom the field · not in the guide
| - Johari, R., Koomen, P., Pekelis, L., & Walsh, D. (2022). Always valid inference: Continuous monitoring of A/B tests. Operations Research, 70(3), 1806–1821. doi ↗
- Wald, A. (1947). Sequential analysis. John Wiley & Sons. doi ↗
|
|---|
| Stratified randomisation | Random assignment carried out separately within groups defined by a key characteristic, such as school or prior attainment, so that the arms are balanced on it. | - R13Experimentation and product analyticsFrom the field · not in the guide
| - Torgerson, D. J., & Torgerson, C. J. (2008). Designing randomised trials in health, education and the social sciences: An introduction. Palgrave Macmillan. doi ↗
|
|---|
| Subgroup analysis | Estimating the effect separately for parts of the sample, such as disadvantaged pupils. Subgroups chosen after seeing the data produce many false positives. | - R13Experimentation and product analyticsFrom the field · not in the guide
| - Education Endowment Foundation. (2022). Statistical analysis guidance for EEF evaluations. d2tic4wvo1iusb.cloudfront.net ↗
- Kohavi, R., Tang, D., & Xu, Y. (2020). Trustworthy online controlled experiments: A practical guide to A/B testing. Cambridge University Press. doi ↗
|
|---|