Gerald Ajam

Library / Glossary / The craft glossary

Basic statistics for educational technologists glossary

50 terms from R18, Basic statistics for educational technologists — 25 defined in the guide itself and 25 more from the field around it. Every term the guide teaches links to the slide that teaches it.

The whole craft glossary

R18 · How to measure & scale

Basic statistics for educational technologists

What can a number tell you, and how far should you trust it?

Every term below is defined in the words of basic statistics for educational technologists, guide R18 of craft guides for educational technologists, and opens the guide at the slide where it is taught. 25 of the 50 are the field’s vocabulary rather than the guide’s own: words a reader will meet around this subject, defined here because the guide assumes them. 4 terms are also defined by another guide in the series; where the two differ, both wordings are given. The whole craft glossary holds all of them together.

TermDefinitionReferred to inRead further
A
Absolute and relative risk

Absolute risk is the chance of an outcome. Relative risk compares two chances as a ratio, so 'twice as likely' can describe a rise from 1 to 2 in 1,000.

  • R18Basic statistics for educational technologistsFrom the field · not in the guide
  • Gigerenzer, G., Gaissmaier, W., Kurz-Milcke, E., Schwartz, L. M., & Woloshin, S. (2007). Helping doctors and patients make sense of health statistics. Psychological Science in the Public Interest, 8(2), 53–96. doi
  • Spiegelhalter, D. (2019). The art of statistics: Learning from data. Pelican.
Analysis of variance (ANOVA)

A test of whether the means of three or more groups differ, made by comparing the variation between groups with the variation inside them.

  • R18Basic statistics for educational technologistsFrom the field · not in the guide
  • Fisher, R. A. (1925). Statistical methods for research workers. Oliver and Boyd.
  • Field, A. (2018). Discovering statistics using IBM SPSS statistics (5th ed.). SAGE.
B
Base rate

How common an outcome is in the population before any test or model is applied.

R15 How common an outcome is in a group before any prediction is made.

  • Bowers, A. J., Sprott, R., & Taff, S. A. (2013). Do we know who will drop out? A review of the predictors of dropping out of high school: Precision, sensitivity, and specificity. The High School Journal, 96(2), 77–100. doi
  • Gigerenzer, G., Gaissmaier, W., Kurz-Milcke, E., Schwartz, L. M., & Woloshin, S. (2007). Helping doctors and patients make sense of health statistics. Psychological Science in the Public Interest, 8(2), 53–96. doi
  • Meehl, P. E., & Rosen, A. (1955). Antecedent probability and the efficiency of psychometric signs, patterns, or cutting scores. Psychological Bulletin, 52(3), 194–216. doi
Bayes' theorem

The rule for updating the probability of something in the light of new evidence, combining how likely it was beforehand with how well the evidence fits.

  • R18Basic statistics for educational technologistsFrom the field · not in the guide
  • Spiegelhalter, D. (2019). The art of statistics: Learning from data. Pelican.
  • Gigerenzer, G., & Hoffrage, U. (1995). How to improve Bayesian reasoning without instruction: Frequency formats. Psychological Review, 102(4), 684–704. doi
  • Gelman, A., Carlin, J. B., Stern, H. S., Dunson, D. B., Vehtari, A., & Rubin, D. B. (2013). Bayesian data analysis (3rd ed.). CRC Press. doi
C
Cohen's d

A standardised effect size that divides the difference between groups by the spread, so studies that use different tests can be compared (Cohen, 1988).

  • Cohen, J. (1988). Statistical power analysis for the behavioral sciences (2nd ed.). Lawrence Erlbaum Associates.
  • Kraft, M. A. (2020). Interpreting effect sizes of education interventions. Educational Researcher, 49(4), 241–253. doi
Cohen's kappa

A measure of how far two raters agree when sorting items into categories, corrected for the agreement expected by chance.

  • R18Basic statistics for educational technologistsFrom the field · not in the guide
  • Cohen, J. (1960). A coefficient of agreement for nominal scales. Educational and Psychological Measurement, 20(1), 37–46. doi
  • Landis, J. R., & Koch, G. G. (1977). The measurement of observer agreement for categorical data. Biometrics, 33(1), 159–174. doi
Confidence interval

A range of plausible values for a true figure, built by a method that captures it a stated share of the time.

  • Cumming, G. (2014). The new statistics: Why and how. Psychological Science, 25(1), 7–29. doi
  • Greenland, S., Senn, S. J., Rothman, K. J., Carlin, J. B., Poole, C., Goodman, S. N., & Altman, D. G. (2016). Statistical tests, P values, confidence intervals, and power: A guide to misinterpretations. European Journal of Epidemiology, 31(4), 337–350. doi
Confounder

A factor that influences both a supposed cause and its outcome, creating a misleading association.

  • Pearl, J., & Mackenzie, D. (2018). The book of why: The new science of cause and effect. Basic Books.
  • Angrist, J. D., & Pischke, J.-S. (2009). Mostly harmless econometrics: An empiricist's companion. Princeton University Press.
Correlation coefficient

A number from −1 to +1 that shows how closely two variables move together in a straight line. Usually written r, it says nothing about what causes what.

  • R18Basic statistics for educational technologistsFrom the field · not in the guide
  • Spiegelhalter, D. (2019). The art of statistics: Learning from data. Pelican.
  • Rodgers, J. L., & Nicewander, W. A. (1988). Thirteen ways to look at the correlation coefficient. The American Statistician, 42(1), 59–66. doi
Counterfactual

What would have happened without the intervention. Never observed directly.

R13 What would have happened to the same people without the change. Never observed; estimated by a control group.

  • Kohavi, R., Tang, D., & Xu, Y. (2020). Trustworthy online controlled experiments: A practical guide to A/B testing. Cambridge University Press. doi
  • Pearl, J., & Mackenzie, D. (2018). The book of why: The new science of cause and effect. Basic Books.
  • Rubin, D. B. (1974). Estimating causal effects of treatments in randomized and nonrandomized studies. Journal of Educational Psychology, 66(5), 688–701. doi
Cronbach's alpha

A figure from 0 to 1 showing how consistently the items in a test or questionnaire measure the same thing. Values of about 0.7 and above are usually accepted.

  • R18Basic statistics for educational technologistsFrom the field · not in the guide
  • Cronbach, L. J. (1951). Coefficient alpha and the internal structure of tests. Psychometrika, 16(3), 297–334. doi
  • Nunnally, J. C. (1978). Psychometric theory (2nd ed.). McGraw-Hill.
D
Distribution

The pattern of how values spread out: where they cluster, how far they range and which way they lean.

  • Spiegelhalter, D. (2019). The art of statistics: Learning from data. Pelican.
  • Kunin, D., Guo, J., Devlin, T. D., & Xiang, D. (n.d.). Seeing theory: A visual introduction to probability and statistics. Brown University. seeing-theory.brown.edu
E
Effect size

The size of a difference or relationship, in raw units or standardised against the spread.

R16 The size of a difference between groups, often in standard deviations. Comparable only when tests and comparison groups are similar.

R21 A difference between groups expressed in standard deviations, so results from different tests can be compared.

  • von Hippel, P. T. (2024). Two-sigma tutoring: Separating science fiction from science fact. Education Next, 24(2). educationnext.org
  • Cohen, J. (1988). Statistical power analysis for the behavioral sciences (2nd ed.). Lawrence Erlbaum Associates.
  • Kraft, M. A., Blazar, D., & Hogan, D. (2018). The effect of teacher coaching on instruction and achievement: A meta-analysis of the causal evidence. Review of Educational Research, 88(4), 547–588. doi
G
Goodhart’s law

When a measure becomes a target, people optimise the measure: logins rise and learning doesn't. Also known as Campbell's law, after the same point made about social indicators (Campbell, 1979).

R15 The observation that a measure used as a target stops being a good measure, in Strathern’s (1997) wording.

  • Strathern, M. (1997). ‘Improving ratings’: Audit in the British university system. European Review, 5(3), 305–321. cambridge.org
  • Campbell, D. T. (1979). Assessing the impact of planned social change. Evaluation and Program Planning, 2(1), 67–90. doi
I
Interquartile range (IQR)

The distance between the 25th and 75th percentiles, which holds the middle half of the values. A measure of spread that outliers do not distort.

  • R18Basic statistics for educational technologistsFrom the field · not in the guide
  • Spiegelhalter, D. (2019). The art of statistics: Learning from data. Pelican.
  • Tukey, J. W. (1977). Exploratory data analysis. Addison-Wesley. doi
L
Linear regression

A method that fits a straight-line relationship between an outcome and one or more predictors, estimating how much the outcome changes as each predictor changes.

  • R18Basic statistics for educational technologistsFrom the field · not in the guide
  • Angrist, J. D., & Pischke, J.-S. (2009). Mostly harmless econometrics: An empiricist's companion. Princeton University Press.
  • Hastie, T., Tibshirani, R., & Friedman, J. (2009). The elements of statistical learning: Data mining, inference, and prediction (2nd ed.). Springer.
  • Gelman, A., & Hill, J. (2007). Data analysis using regression and multilevel/hierarchical models. Cambridge University Press. doi
M
Margin of error

The amount added to and subtracted from a survey estimate to give its confidence interval, usually at 95 per cent. It shrinks as the sample grows.

  • R18Basic statistics for educational technologistsFrom the field · not in the guide
  • Spiegelhalter, D. (2019). The art of statistics: Learning from data. Pelican.
  • Cumming, G. (2014). The new statistics: Why and how. Psychological Science, 25(1), 7–29. doi
Mean

The sum of a set of values divided by how many there are. It is what most people call the average, and a few extreme values can pull it a long way.

  • R18Basic statistics for educational technologistsFrom the field · not in the guide
  • Spiegelhalter, D. (2019). The art of statistics: Learning from data. Pelican.
  • Wheelan, C. (2013). Naked statistics: Stripping the dread from the data. W. W. Norton.
Median

The middle value when a set of values is put in order. Half lie above it and half below, so extreme values barely move it.

  • R18Basic statistics for educational technologistsFrom the field · not in the guide
  • Spiegelhalter, D. (2019). The art of statistics: Learning from data. Pelican.
  • Wheelan, C. (2013). Naked statistics: Stripping the dread from the data. W. W. Norton.
Meta-analysis

A statistical method that pools the effect sizes from many studies of the same question into one weighted average.

  • R18Basic statistics for educational technologistsFrom the field · not in the guide
  • Hattie, J. (2009). Visible learning: A synthesis of over 800 meta-analyses relating to achievement. Routledge.
  • Glass, G. V. (1976). Primary, secondary, and meta-analysis of research. Educational Researcher, 5(10), 3–8. doi
  • Borenstein, M., Hedges, L. V., Higgins, J. P. T., & Rothstein, H. R. (2009). Introduction to meta-analysis. Wiley. doi
Missing data

Values that were never recorded, such as pupils absent on test day. If they are missing for a reason linked to the outcome, analysing only complete cases biases the result.

  • R18Basic statistics for educational technologistsFrom the field · not in the guide
  • Rubin, D. B. (1976). Inference and missing data. Biometrika, 63(3), 581–592. doi
  • Little, R. J. A., & Rubin, D. B. (2002). Statistical analysis with missing data (2nd ed.). Wiley.
Multiple comparisons

The problem that running many tests makes it likely some will come out significant by chance alone. Corrections such as Bonferroni's raise the bar for each test.

  • R18Basic statistics for educational technologistsFrom the field · not in the guide
  • Spiegelhalter, D. (2019). The art of statistics: Learning from data. Pelican.
  • Benjamini, Y., & Hochberg, Y. (1995). Controlling the false discovery rate: A practical and powerful approach to multiple testing. Journal of the Royal Statistical Society: Series B (Methodological), 57(1), 289–300. doi
N
Normal distribution

The symmetrical, bell-shaped distribution in which most values sit near the mean and extreme values are rare. Many common tests assume it.

  • R18Basic statistics for educational technologistsFrom the field · not in the guide
  • Spiegelhalter, D. (2019). The art of statistics: Learning from data. Pelican.
  • Wheelan, C. (2013). Naked statistics: Stripping the dread from the data. W. W. Norton.
Null hypothesis

The assumption of no effect that a significance test measures data against.

  • Wasserstein, R. L., & Lazar, N. A. (2016). The ASA statement on p-values: Context, process, and purpose. The American Statistician, 70(2), 129–133. doi
  • Greenland, S., Senn, S. J., Rothman, K. J., Carlin, J. B., Poole, C., Goodman, S. N., & Altman, D. G. (2016). Statistical tests, P values, confidence intervals, and power: A guide to misinterpretations. European Journal of Epidemiology, 31(4), 337–350. doi
O
Outlier

A value far from the rest of the data. It may be an error or a real and important case, so it should be investigated before it is removed.

  • R18Basic statistics for educational technologistsFrom the field · not in the guide
  • Spiegelhalter, D. (2019). The art of statistics: Learning from data. Pelican.
  • Tukey, J. W. (1977). Exploratory data analysis. Addison-Wesley. doi
Overfitting

A model that captures the quirks of past data performs well in-sample and fails on new data. Only out-of-sample performance tests prediction honestly (Hastie et al., 2009).

  • Hastie, T., Tibshirani, R., & Friedman, J. (2009). The elements of statistical learning: Data mining, inference, and prediction (2nd ed.). Springer.
P
p-hacking

Running many analyses on the same data, which raises the odds that one crosses 0.05 by chance (Simmons et al., 2011).

  • Simmons, J. P., Nelson, L. D., & Simonsohn, U. (2011). False-positive psychology: Undisclosed flexibility in data collection and analysis allows presenting anything as significant. Psychological Science, 22(11), 1359–1366. doi
p-value

The probability of data at least this extreme if there were no real effect.

  • Wasserstein, R. L., & Lazar, N. A. (2016). The ASA statement on p-values: Context, process, and purpose. The American Statistician, 70(2), 129–133. doi
  • Greenland, S., Senn, S. J., Rothman, K. J., Carlin, J. B., Poole, C., Goodman, S. N., & Altman, D. G. (2016). Statistical tests, P values, confidence intervals, and power: A guide to misinterpretations. European Journal of Epidemiology, 31(4), 337–350. doi
  • Simmons, J. P., Nelson, L. D., & Simonsohn, U. (2011). False-positive psychology: Undisclosed flexibility in data collection and analysis allows presenting anything as significant. Psychological Science, 22(11), 1359–1366. doi
Percentile

The value below which a given percentage of results fall. A student at the 80th percentile scored higher than 80 per cent of the group.

  • R18Basic statistics for educational technologistsFrom the field · not in the guide
  • Wheelan, C. (2013). Naked statistics: Stripping the dread from the data. W. W. Norton.
Population

The whole group a conclusion is meant to apply to.

  • Spiegelhalter, D. (2019). The art of statistics: Learning from data. Pelican.
Power

The probability a study detects an effect of a given size if it truly exists.

  • Cohen, J. (1988). Statistical power analysis for the behavioral sciences (2nd ed.). Lawrence Erlbaum Associates.
  • Greenland, S., Senn, S. J., Rothman, K. J., Carlin, J. B., Poole, C., Goodman, S. N., & Altman, D. G. (2016). Statistical tests, P values, confidence intervals, and power: A guide to misinterpretations. European Journal of Epidemiology, 31(4), 337–350. doi
R
R-squared

The proportion of the variation in an outcome that a regression model accounts for, from 0 to 1. A high value does not show that the model is causal.

  • R18Basic statistics for educational technologistsFrom the field · not in the guide
  • Gelman, A., & Hill, J. (2007). Data analysis using regression and multilevel/hierarchical models. Cambridge University Press. doi
  • Field, A. (2018). Discovering statistics using IBM SPSS statistics (5th ed.). SAGE.
Randomisation

Assigning units to groups by chance, so groups are balanced on measured and unmeasured factors.

  • Angrist, J. D., & Pischke, J.-S. (2009). Mostly harmless econometrics: An empiricist's companion. Princeton University Press.
  • Pearl, J., & Mackenzie, D. (2018). The book of why: The new science of cause and effect. Basic Books.
Regression discontinuity

A design that compares students just above and just below a cutoff, such as a scholarship score, who are otherwise comparable. Only valid near the cutoff (Thistlethwaite & Campbell, 1960).

  • Thistlethwaite, D. L., & Campbell, D. T. (1960). Regression-discontinuity analysis: An alternative to the ex post facto experiment. Journal of Educational Psychology, 51(6), 309–317. doi
  • Angrist, J. D., & Pischke, J.-S. (2009). Mostly harmless econometrics: An empiricist's companion. Princeton University Press.
Regression to the mean

Extreme results are partly luck, and luck doesn't repeat, so schools chosen because they scored lowest will tend to score higher next time, with or without help.

  • Spiegelhalter, D. (2019). The art of statistics: Learning from data. Pelican.
  • Galton, F. (1886). Regression towards mediocrity in hereditary stature. The Journal of the Anthropological Institute of Great Britain and Ireland, 15, 246–263. doi
Reliability

The consistency of a measure across items, occasions or raters.

  • Cronbach, L. J. (1951). Coefficient alpha and the internal structure of tests. Psychometrika, 16(3), 297–334. doi
  • Nunnally, J. C. (1978). Psychometric theory (2nd ed.). McGraw-Hill.
  • Cohen, J. (1960). A coefficient of agreement for nominal scales. Educational and Psychological Measurement, 20(1), 37–46. doi
S
Sample

The subset of a population actually observed.

  • Spiegelhalter, D. (2019). The art of statistics: Learning from data. Pelican.
Selection bias

Distortion that arises when the people in a sample or a comparison group differ systematically from those they are meant to represent, such as schools that volunteered for a pilot.

  • R18Basic statistics for educational technologistsFrom the field · not in the guide
  • Angrist, J. D., & Pischke, J.-S. (2009). Mostly harmless econometrics: An empiricist's companion. Princeton University Press.
  • Heckman, J. J. (1979). Sample selection bias as a specification error. Econometrica, 47(1), 153–161. doi
Sensitivity and specificity

Sensitivity is the share of students who need help that a tool catches; specificity is the share of those who don't that it clears. No threshold removes the trade-off between them.

  • Gigerenzer, G., Gaissmaier, W., Kurz-Milcke, E., Schwartz, L. M., & Woloshin, S. (2007). Helping doctors and patients make sense of health statistics. Psychological Science in the Public Interest, 8(2), 53–96. doi
  • Hastie, T., Tibshirani, R., & Friedman, J. (2009). The elements of statistical learning: Data mining, inference, and prediction (2nd ed.). Springer.
Simpson's paradox

A pattern that appears in every subgroup but reverses or vanishes when the subgroups are combined, because the groups differ in size or make-up.

  • R18Basic statistics for educational technologistsFrom the field · not in the guide
  • Charig, C. R., Webb, D. R., Payne, S. R., & Wickham, J. E. A. (1986). Comparison of treatment of renal calculi by open surgery, percutaneous nephrolithotomy, and extracorporeal shockwave lithotripsy. BMJ, 292(6524), 879–882. doi
  • Pearl, J., & Mackenzie, D. (2018). The book of why: The new science of cause and effect. Basic Books.
  • Simpson, E. H. (1951). The interpretation of interaction in contingency tables. Journal of the Royal Statistical Society: Series B (Methodological), 13(2), 238–241. doi
Skew

Lopsidedness in a distribution, with a long tail on one side. In skewed data such as time spent on a platform, the mean and the median can differ widely.

  • R18Basic statistics for educational technologistsFrom the field · not in the guide
  • Spiegelhalter, D. (2019). The art of statistics: Learning from data. Pelican.
Standard deviation

Roughly the typical distance of values from their mean.

  • Spiegelhalter, D. (2019). The art of statistics: Learning from data. Pelican.
  • Kunin, D., Guo, J., Devlin, T. D., & Xiang, D. (n.d.). Seeing theory: A visual introduction to probability and statistics. Brown University. seeing-theory.brown.edu
Standard error

The typical wobble of an estimate across repeated samples. Shrinks with the square root of sample size.

  • Cumming, G. (2014). The new statistics: Why and how. Psychological Science, 25(1), 7–29. doi
  • Spiegelhalter, D. (2019). The art of statistics: Learning from data. Pelican.
Standard error of measurement

An estimate of how far an observed score might sit from a student's true score, which makes a single score a range, not a point.

  • Nunnally, J. C. (1978). Psychometric theory (2nd ed.). McGraw-Hill.
Statistical significance

A p-value below a chosen threshold, conventionally 0.05. A convention, not a verdict.

  • Wasserstein, R. L., & Lazar, N. A. (2016). The ASA statement on p-values: Context, process, and purpose. The American Statistician, 70(2), 129–133. doi
  • Greenland, S., Senn, S. J., Rothman, K. J., Carlin, J. B., Poole, C., Goodman, S. N., & Altman, D. G. (2016). Statistical tests, P values, confidence intervals, and power: A guide to misinterpretations. European Journal of Epidemiology, 31(4), 337–350. doi
  • Simmons, J. P., Nelson, L. D., & Simonsohn, U. (2011). False-positive psychology: Undisclosed flexibility in data collection and analysis allows presenting anything as significant. Psychological Science, 22(11), 1359–1366. doi
T
t-test

A test of whether the difference between two means is larger than chance variation alone would plausibly produce.

  • R18Basic statistics for educational technologistsFrom the field · not in the guide
  • Student. (1908). The probable error of a mean. Biometrika, 6(1), 1–25. doi
  • Field, A. (2018). Discovering statistics using IBM SPSS statistics (5th ed.). SAGE.
Type I and Type II errors

A Type I error is concluding there is an effect when there is none. A Type II error is failing to detect an effect that is real.

  • R18Basic statistics for educational technologistsFrom the field · not in the guide
  • Cohen, J. (1988). Statistical power analysis for the behavioral sciences (2nd ed.). Lawrence Erlbaum Associates.
  • Greenland, S., Senn, S. J., Rothman, K. J., Carlin, J. B., Poole, C., Goodman, S. N., & Altman, D. G. (2016). Statistical tests, P values, confidence intervals, and power: A guide to misinterpretations. European Journal of Epidemiology, 31(4), 337–350. doi
  • Neyman, J., & Pearson, E. S. (1933). On the problem of the most efficient tests of statistical hypotheses. Philosophical Transactions of the Royal Society of London. Series A, Containing Papers of a Mathematical or Physical Character, 231, 289–337. doi
V
Validity

The degree to which evidence supports the intended meaning and use of scores.

  • Messick, S. (1995). Validity of psychological assessment: Validation of inferences from persons' responses and performances as scientific inquiry into score meaning. American Psychologist, 50(9), 741–749. doi
  • Nunnally, J. C. (1978). Psychometric theory (2nd ed.). McGraw-Hill.
Variance

The average squared distance of values from their mean. The standard deviation is its square root, and 'variance explained' describes how much of it a model accounts for.

  • R18Basic statistics for educational technologistsFrom the field · not in the guide
  • Spiegelhalter, D. (2019). The art of statistics: Learning from data. Pelican.
  • Wheelan, C. (2013). Naked statistics: Stripping the dread from the data. W. W. Norton.
  • Field, A. (2018). Discovering statistics using IBM SPSS statistics (5th ed.). SAGE.
Z
z-score

A value expressed as the number of standard deviations it lies above or below the mean, which lets results on different scales be compared.

  • R18Basic statistics for educational technologistsFrom the field · not in the guide
  • Wheelan, C. (2013). Naked statistics: Stripping the dread from the data. W. W. Norton.
  • Field, A. (2018). Discovering statistics using IBM SPSS statistics (5th ed.). SAGE.
Singapore