Gerald Ajam

Library / Glossary / The craft glossary

Experimentation and product analytics glossary

45 terms from R13, Experimentation and product analytics — 21 defined in the guide itself and 24 more from the field around it. Every term the guide teaches links to the slide that teaches it.

The whole craft glossary

R13 · How to build

Experimentation and product analytics

Did this change do what we hoped?

Every term below is defined in the words of experimentation and product analytics, guide R13 of craft guides for educational technologists, and opens the guide at the slide where it is taught. 24 of the 45 are the field’s vocabulary rather than the guide’s own: words a reader will meet around this subject, defined here because the guide assumes them. 2 terms are also defined by another guide in the series; where the two differ, both wordings are given. The whole craft glossary holds all of them together.

TermDefinitionReferred to inRead further
A
A/A test

A test in which both groups get the same version. It checks the system: any "significant" result is a false positive.

  • Kohavi, R., Tang, D., & Xu, Y. (2020). Trustworthy online controlled experiments: A practical guide to A/B testing. Cambridge University Press. doi
A/B test

An experiment in which users are randomly split between the current version (A) and a changed version (B), and the groups' outcomes are compared to estimate the effect of the change.

  • R13Experimentation and product analyticsFrom the field · not in the guide
  • Kohavi, R., Tang, D., & Xu, Y. (2020). Trustworthy online controlled experiments: A practical guide to A/B testing. Cambridge University Press. doi
Always-valid p-value

A p-value designed to stay correct however often a running test is checked (Johari et al., 2022).

  • Johari, R., Koomen, P., Pekelis, L., & Walsh, D. (2022). Always valid inference: Continuous monitoring of A/B tests. Operations Research, 70(3), 1806–1821. doi
Assent

Agreement to take part given by children too young to consent; principles of consent apply to children and young people as well as to adults (British Educational Research Association, 2024).

R04 A child's own agreement to take part in research, sought alongside a guardian's consent (British Educational Research Association, 2024).

R05 A child's own agreement to take part, sought alongside consent from a responsible adult (British Educational Research Association, 2024).

  • British Educational Research Association. (2024). Ethical guidelines for educational research (5th ed.). bera.ac.uk
  • Lundy, L. (2007). ‘Voice’ is not enough: Conceptualising Article 12 of the United Nations Convention on the Rights of the Child. British Educational Research Journal, 33(6), 927–942. doi
  • United Nations. (1989). Convention on the Rights of the Child (General Assembly resolution 44/25). ohchr.org
Attrition

The loss of participants from a study before outcomes are measured. Differential attrition, where one group loses more than the other, can bias the estimated effect.

  • R13Experimentation and product analyticsFrom the field · not in the guide
  • What Works Clearinghouse. (2022). What Works Clearinghouse procedures and standards handbook, version 5.0 (WWC 2022008). U.S. Department of Education, Institute of Education Sciences. ies.ed.gov
  • Education Endowment Foundation. (2022). Statistical analysis guidance for EEF evaluations. d2tic4wvo1iusb.cloudfront.net
B
Baseline equivalence

Evidence that the intervention and comparison groups were similar on key characteristics, such as prior attainment, before the intervention began.

  • R13Experimentation and product analyticsFrom the field · not in the guide
  • What Works Clearinghouse. (2022). What Works Clearinghouse procedures and standards handbook, version 5.0 (WWC 2022008). U.S. Department of Education, Institute of Education Sciences. ies.ed.gov
Before-and-after comparison

Comparing a measure before and after a change. It credits the change with everything else that happened in between, such as a new term, an exam season or a new timetable.

  • Kohavi, R., Tang, D., & Xu, Y. (2020). Trustworthy online controlled experiments: A practical guide to A/B testing. Cambridge University Press. doi
C
Cluster

A group whose members share an experience and tend to resemble each other: a class, a teacher's classes or a school.

  • Donner, A., & Klar, N. (2000). Design and analysis of cluster randomization trials in health research. Arnold. wiley.com
  • Hedges, L. V., & Hedberg, E. C. (2007). Intraclass correlation values for planning group-randomized trials in education. Educational Evaluation and Policy Analysis, 29(1), 60–87. doi
  • Education Endowment Foundation. (2022). Statistical analysis guidance for EEF evaluations. d2tic4wvo1iusb.cloudfront.net
Cluster randomised trial

A trial that randomises whole groups, such as classes or schools, rather than individual pupils. It needs a larger sample than an individually randomised trial of the same precision.

  • R13Experimentation and product analyticsFrom the field · not in the guide
  • Donner, A., & Klar, N. (2000). Design and analysis of cluster randomization trials in health research. Arnold. wiley.com
  • Education Endowment Foundation. (2022). Statistical analysis guidance for EEF evaluations. d2tic4wvo1iusb.cloudfront.net
Cohort analysis

Analysis that groups users by a shared starting point, such as the week they joined, and follows each group over time, so that changes in behaviour are not hidden by a changing user mix.

  • R13Experimentation and product analyticsFrom the field · not in the guide
  • Croll, A., & Yoskovitz, B. (2013). Lean analytics: Use data to build a better startup faster. O'Reilly Media.
Complier average causal effect (CACE)

An estimate of the effect of an intervention on those who took it up as intended, reported alongside the intention-to-treat estimate.

  • R13Experimentation and product analyticsFrom the field · not in the guide
  • Education Endowment Foundation. (2022). Statistical analysis guidance for EEF evaluations. d2tic4wvo1iusb.cloudfront.net
  • Angrist, J. D., Imbens, G. W., & Rubin, D. B. (1996). Identification of causal effects using instrumental variables. Journal of the American Statistical Association, 91(434), 444–455. doi
Contamination

Spillover of one version's effects into the other group, for example through shared answers or a shared lesson.

  • Donner, A., & Klar, N. (2000). Design and analysis of cluster randomization trials in health research. Arnold. wiley.com
  • Ritter, S., Murphy, A., Fancsali, S. E., Fitkariwala, V., Patel, N., & Lomas, J. D. (2020). UpGrade: An open source tool to support A/B testing in educational software [Paper presentation]. Workshop on Educational A/B Testing at Scale, Learning @ Scale 2020. upgradeplatform.org
Control group

A group like the one that got the change, in the same weeks, that did not get it. It shows how much of a rise would have happened anyway.

  • Kohavi, R., Tang, D., & Xu, Y. (2020). Trustworthy online controlled experiments: A practical guide to A/B testing. Cambridge University Press. doi
  • What Works Clearinghouse. (2022). What Works Clearinghouse procedures and standards handbook, version 5.0 (WWC 2022008). U.S. Department of Education, Institute of Education Sciences. ies.ed.gov
Counterfactual

What would have happened to the same people without the change. Never observed; estimated by a control group.

R18 What would have happened without the intervention. Never observed directly.

  • Kohavi, R., Tang, D., & Xu, Y. (2020). Trustworthy online controlled experiments: A practical guide to A/B testing. Cambridge University Press. doi
  • Pearl, J., & Mackenzie, D. (2018). The book of why: The new science of cause and effect. Basic Books.
  • Rubin, D. B. (1974). Estimating causal effects of treatments in randomized and nonrandomized studies. Journal of Educational Psychology, 66(5), 688–701. doi
CUPED

A variance reduction method that adjusts each user's outcome using data from before the experiment, so that smaller effects can be detected with the same sample.

  • R13Experimentation and product analyticsFrom the field · not in the guide
  • Kohavi, R., Tang, D., & Xu, Y. (2020). Trustworthy online controlled experiments: A practical guide to A/B testing. Cambridge University Press. doi
  • Deng, A., Xu, Y., Kohavi, R., & Walker, T. (2013). Improving the sensitivity of online controlled experiments by utilizing pre-experiment data. In Proceedings of the Sixth ACM International Conference on Web Search and Data Mining (pp. 123–132). ACM. doi
D
Design effect

The factor by which clustering inflates the sample needed: 1 + (m − 1) × ICC (Donner & Klar, 2000).

  • Kish, L. (1965). Survey sampling. Wiley. wiley.com
  • Donner, A., & Klar, N. (2000). Design and analysis of cluster randomization trials in health research. Arnold. wiley.com
  • Hedges, L. V., & Hedberg, E. C. (2007). Intraclass correlation values for planning group-randomized trials in education. Educational Evaluation and Policy Analysis, 29(1), 60–87. doi
Difference-in-differences

A quasi-experimental method that compares the change over time in a group that received an intervention with the change in a group that did not.

  • R13Experimentation and product analyticsFrom the field · not in the guide
  • Angrist, J. D., & Pischke, J.-S. (2009). Mostly harmless econometrics: An empiricist's companion. Princeton University Press. doi
E
External validity

The extent to which a study's result holds for other people, settings, times and versions of the intervention than those studied.

  • R13Experimentation and product analyticsFrom the field · not in the guide
  • Shadish, W. R., Cook, T. D., & Campbell, D. T. (2002). Experimental and quasi-experimental designs for generalized causal inference. Houghton Mifflin.
G
Guardrail metric

A measure that must not get worse during a test, such as teacher workload or results for the lowest attainers.

  • Kohavi, R., Tang, D., & Xu, Y. (2020). Trustworthy online controlled experiments: A practical guide to A/B testing. Cambridge University Press. doi
H
Holdout group

A small share of users kept on the old experience for a long period, often across many launches, to measure the long-term or cumulative effect of changes.

  • R13Experimentation and product analyticsFrom the field · not in the guide
  • Kohavi, R., Tang, D., & Xu, Y. (2020). Trustworthy online controlled experiments: A practical guide to A/B testing. Cambridge University Press. doi
I
Institutional review board (IRB)

A committee that reviews planned research with human participants to protect their rights and welfare. Called a research ethics committee in the UK.

  • R13Experimentation and product analyticsFrom the field · not in the guide
  • Protection of Human Subjects, 45 C.F.R. pt. 46 (2018). ecfr.gov
  • British Educational Research Association. (2024). Ethical guidelines for educational research (5th ed.). bera.ac.uk
Instrumentation

The code that records user and system events in a product, and the work of adding it. An experiment can only measure what has been instrumented.

  • R13Experimentation and product analyticsFrom the field · not in the guide
  • Kohavi, R., Tang, D., & Xu, Y. (2020). Trustworthy online controlled experiments: A practical guide to A/B testing. Cambridge University Press. doi
Intention to treat

Analysing everyone by the group they were assigned to, whatever they actually used (Education Endowment Foundation, 2022).

  • Education Endowment Foundation. (2022). Statistical analysis guidance for EEF evaluations. d2tic4wvo1iusb.cloudfront.net
  • What Works Clearinghouse. (2022). What Works Clearinghouse procedures and standards handbook, version 5.0 (WWC 2022008). U.S. Department of Education, Institute of Education Sciences. ies.ed.gov
Internal validity

The extent to which a study supports the claim that the intervention, and not something else, caused the observed difference.

  • R13Experimentation and product analyticsFrom the field · not in the guide
  • Shadish, W. R., Cook, T. D., & Campbell, D. T. (2002). Experimental and quasi-experimental designs for generalized causal inference. Houghton Mifflin.
Interrupted time series

A quasi-experimental design that takes many measurements before and after an intervention and looks for a change in level or trend at the point it was introduced.

  • R13Experimentation and product analyticsFrom the field · not in the guide
  • Shadish, W. R., Cook, T. D., & Campbell, D. T. (2002). Experimental and quasi-experimental designs for generalized causal inference. Houghton Mifflin.
Intraclass correlation (ICC)

The share of the variation in an outcome that lies between clusters rather than within them.

  • Hedges, L. V., & Hedberg, E. C. (2007). Intraclass correlation values for planning group-randomized trials in education. Educational Evaluation and Policy Analysis, 29(1), 60–87. doi
  • Donner, A., & Klar, N. (2000). Design and analysis of cluster randomization trials in health research. Arnold. wiley.com
M
Minimum detectable effect

The smallest true effect a test is likely to detect, given its size, design and power.

  • Hedges, L. V., & Hedberg, E. C. (2007). Intraclass correlation values for planning group-randomized trials in education. Educational Evaluation and Policy Analysis, 29(1), 60–87. doi
  • Bloom, H. S. (1995). Minimum detectable effects: A simple way to report the statistical power of experimental designs. Evaluation Review, 19(5), 547–556. doi
Multi-armed bandit

An allocation method that shifts traffic towards the better-performing versions while the test runs, trading some certainty about the effect for fewer users on a worse version.

  • R13Experimentation and product analyticsFrom the field · not in the guide
  • Lattimore, T., & Szepesvári, C. (2020). Bandit algorithms. Cambridge University Press. doi
N
North star

A product team's long-term objective. In a learning product, usually a learning outcome that takes too long to measure in one test.

  • Kohavi, R., Tang, D., & Xu, Y. (2020). Trustworthy online controlled experiments: A practical guide to A/B testing. Cambridge University Press. doi
Novelty effect

A temporary change in behaviour because something is new, which fades as users get used to it (Kohavi et al., 2012).

  • Kohavi, R., Deng, A., Frasca, B., Longbotham, R., Walker, T., & Xu, Y. (2012). Trustworthy online controlled experiments: Five puzzling outcomes explained. In Proceedings of the 18th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (pp. 786–794). Association for Computing Machinery. doi
  • Clark, R. E. (1983). Reconsidering research on learning from media. Review of Educational Research, 53(4), 445–459. doi
  • Kohavi, R., Tang, D., & Xu, Y. (2020). Trustworthy online controlled experiments: A practical guide to A/B testing. Cambridge University Press. doi
O
Optional stopping

Checking a running test repeatedly and stopping when the result looks good. It inflates false positives.

  • Johari, R., Koomen, P., Pekelis, L., & Walsh, D. (2022). Always valid inference: Continuous monitoring of A/B tests. Operations Research, 70(3), 1806–1821. doi
  • Simmons, J. P., Nelson, L. D., & Simonsohn, U. (2011). False-positive psychology: Undisclosed flexibility in data collection and analysis allows presenting anything as significant. Psychological Science, 22(11), 1359–1366. doi
  • Kohavi, R., Tang, D., & Xu, Y. (2020). Trustworthy online controlled experiments: A practical guide to A/B testing. Cambridge University Press. doi
Overall evaluation criterion

The primary measure a test is designed to move, chosen in advance (Kohavi et al., 2020).

  • Kohavi, R., Tang, D., & Xu, Y. (2020). Trustworthy online controlled experiments: A practical guide to A/B testing. Cambridge University Press. doi
  • Kohavi, R., Deng, A., Frasca, B., Longbotham, R., Walker, T., & Xu, Y. (2012). Trustworthy online controlled experiments: Five puzzling outcomes explained. In Proceedings of the 18th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (pp. 786–794). Association for Computing Machinery. doi
P
Pre-registration

A dated record of a test's outcome, analysis and decision rule, made before the data arrives.

  • Simmons, J. P., Nelson, L. D., & Simonsohn, U. (2011). False-positive psychology: Undisclosed flexibility in data collection and analysis allows presenting anything as significant. Psychological Science, 22(11), 1359–1366. doi
  • Education Endowment Foundation. (2022). Statistical analysis guidance for EEF evaluations. d2tic4wvo1iusb.cloudfront.net
  • Inter-university Consortium for Political and Social Research. (n.d.). Registry of Efficacy and Effectiveness Studies (REES). Retrieved September 17, 2026, from icpsr.umich.edu
Primacy effect

A short-lived dip in a metric after a change, because experienced users need time to adjust to it. The opposite of the novelty effect.

  • R13Experimentation and product analyticsFrom the field · not in the guide
  • Kohavi, R., Tang, D., & Xu, Y. (2020). Trustworthy online controlled experiments: A practical guide to A/B testing. Cambridge University Press. doi
Propensity score matching

A method that pairs participants and non-participants with a similar estimated probability of taking part, given their observed characteristics, to make the groups more comparable.

  • R13Experimentation and product analyticsFrom the field · not in the guide
  • Rosenbaum, P. R., & Rubin, D. B. (1983). The central role of the propensity score in observational studies for causal effects. Biometrika, 70(1), 41–55. doi
Q
Quasi-experiment

A study that estimates the effect of an intervention without random assignment, using a comparison group or a time pattern chosen to rule out other explanations.

  • R13Experimentation and product analyticsFrom the field · not in the guide
  • What Works Clearinghouse. (2022). What Works Clearinghouse procedures and standards handbook, version 5.0 (WWC 2022008). U.S. Department of Education, Institute of Education Sciences. ies.ed.gov
  • Shadish, W. R., Cook, T. D., & Campbell, D. T. (2002). Experimental and quasi-experimental designs for generalized causal inference. Houghton Mifflin.
R
Randomisation unit

The thing chance assigns to a version: a student, a class, a teacher or a school.

  • Kohavi, R., Tang, D., & Xu, Y. (2020). Trustworthy online controlled experiments: A practical guide to A/B testing. Cambridge University Press. doi
  • Ritter, S., Murphy, A., Fancsali, S. E., Fitkariwala, V., Patel, N., & Lomas, J. D. (2020). UpGrade: An open source tool to support A/B testing in educational software [Paper presentation]. Workshop on Educational A/B Testing at Scale, Learning @ Scale 2020. upgradeplatform.org
  • What Works Clearinghouse. (2022). What Works Clearinghouse procedures and standards handbook, version 5.0 (WWC 2022008). U.S. Department of Education, Institute of Education Sciences. ies.ed.gov
Randomised controlled trial (RCT)

A study that assigns participants at random to an intervention or a control condition, so that a difference in outcomes can be attributed to the intervention.

  • R13Experimentation and product analyticsFrom the field · not in the guide
  • What Works Clearinghouse. (2022). What Works Clearinghouse procedures and standards handbook, version 5.0 (WWC 2022008). U.S. Department of Education, Institute of Education Sciences. ies.ed.gov
  • Torgerson, D. J., & Torgerson, C. J. (2008). Designing randomised trials in health, education and the social sciences: An introduction. Palgrave Macmillan. doi
Researcher degrees of freedom

The many choices in collecting and analysing data that can push a result towards significance (Simmons et al., 2011).

  • Simmons, J. P., Nelson, L. D., & Simonsohn, U. (2011). False-positive psychology: Undisclosed flexibility in data collection and analysis allows presenting anything as significant. Psychological Science, 22(11), 1359–1366. doi
  • Gelman, A., & Loken, E. (2014). The statistical crisis in science. American Scientist, 102(6), 460–465. doi
S
Sample ratio mismatch (SRM)

A split of users between groups that differs from the planned ratio by more than chance allows. It signals a fault in assignment or logging, so the results should not be trusted.

  • R13Experimentation and product analyticsFrom the field · not in the guide
  • Kohavi, R., Tang, D., & Xu, Y. (2020). Trustworthy online controlled experiments: A practical guide to A/B testing. Cambridge University Press. doi
  • Fabijan, A., Gupchup, J., Gupta, S., Omhover, J., Qin, W., Vermeer, L., & Dmitriev, P. (2019). Diagnosing sample ratio mismatch in online controlled experiments: A taxonomy and rules of thumb for practitioners. In Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining (pp. 2156–2164). ACM. doi
Sequential testing

Methods that allow a running experiment to be analysed at planned or repeated points and stopped early, while keeping the false positive rate at its stated level.

  • R13Experimentation and product analyticsFrom the field · not in the guide
  • Johari, R., Koomen, P., Pekelis, L., & Walsh, D. (2022). Always valid inference: Continuous monitoring of A/B tests. Operations Research, 70(3), 1806–1821. doi
  • Wald, A. (1947). Sequential analysis. John Wiley & Sons. doi
Stratified randomisation

Random assignment carried out separately within groups defined by a key characteristic, such as school or prior attainment, so that the arms are balanced on it.

  • R13Experimentation and product analyticsFrom the field · not in the guide
  • Torgerson, D. J., & Torgerson, C. J. (2008). Designing randomised trials in health, education and the social sciences: An introduction. Palgrave Macmillan. doi
Subgroup analysis

Estimating the effect separately for parts of the sample, such as disadvantaged pupils. Subgroups chosen after seeing the data produce many false positives.

  • R13Experimentation and product analyticsFrom the field · not in the guide
  • Education Endowment Foundation. (2022). Statistical analysis guidance for EEF evaluations. d2tic4wvo1iusb.cloudfront.net
  • Kohavi, R., Tang, D., & Xu, Y. (2020). Trustworthy online controlled experiments: A practical guide to A/B testing. Cambridge University Press. doi
T
Twyman's law

Any figure that looks interesting or different is usually wrong, so the more a result looks like a breakthrough, the more checking it needs (Kohavi et al., 2020).

  • Kohavi, R., Tang, D., & Xu, Y. (2020). Trustworthy online controlled experiments: A practical guide to A/B testing. Cambridge University Press. doi
W
Waitlist control

A control group that receives the intervention after the trial's outcomes have been measured, so that no participant is denied it.

  • R13Experimentation and product analyticsFrom the field · not in the guide
  • Torgerson, D. J., & Torgerson, C. J. (2008). Designing randomised trials in health, education and the social sciences: An introduction. Palgrave Macmillan. doi
Singapore