Gerald Ajam

Library / Glossary / The context glossary

Assessment systems glossary

42 terms from C14, Assessment systems: validity, reliability and high stakes — each defined in the primer's own words. Every term the primer teaches links to the slide that teaches it.

The whole context glossary

C14 · Curriculum and assessment

Assessment systems: validity, reliability and high stakes

“A score measures what the item says it measures.”

Every term below is defined in the words of assessment systems, primer C14 of understanding the context of education, and opens the primer at the slide where it is taught. 1 term is also defined by another primer in the series; where the two differ, both wordings are given. The whole context glossary holds all of them together.

TermDefinitionReferred to inRead further
A
Academic integrity

The expectation that work submitted for assessment is the student's own and that sources and tools are acknowledged. It is a property of the whole assessment design, not only of a rule in a handbook.

  • Weber-Wulff, D., Anohina-Naumeca, A., Bjelobaba, S., Foltýnek, T., Guerrero-Dib, J., Popoola, O., Šigut, P., & Waddington, L. (2023). Testing of detection tools for AI-generated text. International Journal for Educational Integrity, 19, 26. doi
  • Lodge, J. M., Howard, S., Bearman, M., Dawson, P., & Associates. (2023). Assessment reform for the age of artificial intelligence. Tertiary Education Quality and Standards Agency. teqsa.gov.au
  • Liang, W., Yuksekgonul, M., Mao, Y., Wu, E., & Zou, J. (2023). GPT detectors are biased against non-native English writers. Patterns, 4(7), 100779. doi
Achievement Levels

Singapore's primary examination scoring from the 2021 cohort: each subject is reported in one of eight Achievement Levels, and the four subjects sum to a score between 4 and 32 (Ministry of Education, Singapore, 2016).

  • Ministry of Education, Singapore. (2016, July 13). Changes to PSLE scoring and Secondary One posting [Press release]. moe.gov.sg
  • Ministry of Education, Singapore. (n.d.). How will the current PSLE scoring system benefit your child? Retrieved September 18, 2026, from moe.gov.sg
  • Ministry of Education, Singapore. (2025, November 25). Release of 2025 PSLE results [Press release]. moe.gov.sg
AI and assessment design

Redesigning what is assessed, and where, so that the inference from work to student still holds when a capable tool is available. Lodge et al. (2023) argue for authentic engagement with AI plus security at a few meaningful points.

  • Lodge, J. M., Howard, S., Bearman, M., Dawson, P., & Associates. (2023). Assessment reform for the age of artificial intelligence. Tertiary Education Quality and Standards Agency. teqsa.gov.au
  • Bridgeman, A., Weeks, R., & Liu, D. (2024, November 26). Aligning our assessments to the age of generative AI. Teaching@Sydney, University of Sydney. educational-innovation.sydney.edu.au
Assessment literacy

Knowing what assessments can and cannot tell you, and how to build and read them. The term was popularised by Stiggins (1991), who argued that decision-makers need it as much as teachers do.

  • Stiggins, R. J. (1991). Assessment literacy. Phi Delta Kappan, 72(7), 534–539. jstor.org
Audit test

A second measure of the same learning, taken by the same students with no stakes attached; the standard way to detect score inflation, and rarely done.

  • Koretz, D., & Barron, S. I. (1998). The validity of gains in scores on the Kentucky Instructional Results Information System (KIRIS) (MR-1014-EDU). RAND. rand.org
  • Koretz, D. (2008). Measuring up: What educational testing really tells us. Harvard University Press. jstor.org
  • Dee, T. S., & Jacob, B. (2011). The impact of No Child Left Behind on student achievement. Journal of Policy Analysis and Management, 30(3), 418–446. doi
Authentic assessment

Assessment that "directly examines student performance on worthy intellectual tasks" rather than using indirect proxy items (Wiggins, 1990).

  • Wiggins, G. (1990). The case for authentic assessment (ERIC Digest ED328611). ERIC Clearinghouse on Tests, Measurement and Evaluation. files.eric.ed.gov
C
Campbell's law

The more any quantitative social indicator is used for social decision-making, the more subject it will be to corruption pressures and the more apt it will be to distort the processes it monitors (Campbell, 1979).

C01 When test scores become the goal of the teaching process, they both lose their value as indicators of educational status and distort the educational process in undesirable ways (Campbell, 1979).

C17 The more a quantitative social indicator is used for social decision-making, the more subject it is to corruption pressures and the more apt to distort the social processes it is intended to monitor (Campbell, 1979).

  • Campbell, D. T. (1979). Assessing the impact of planned social change. Evaluation and Program Planning, 2(1), 67–90. doi
  • Strathern, M. (1997). ‘Improving ratings’: Audit in the British University system. European Review, 5(3), 305–321. doi
Classical test theory

A model that treats an observed score as the sum of a true score and an independent random error component, giving one error figure for the whole test (AERA et al., 2014).

  • AERA, APA, & NCME. (2014). Standards for educational and psychological testing. American Educational Research Association. testingstandards.net
  • Lord, F. M., & Novick, M. R. (1968). Statistical theories of mental test scores. Addison-Wesley. doi
Comparative judgement

A method in which assessors repeatedly choose the better of two pieces of work, and a scale is built from the pattern of choices rather than from a rubric (Pollitt, 2012).

  • Pollitt, A. (2012). The method of Adaptive Comparative Judgement. Assessment in Education: Principles, Policy & Practice, 19(3), 281–300. doi
  • Bramley, T. (2015). Investigating the reliability of Adaptive Comparative Judgment (Cambridge Assessment Research Report). Cambridge Assessment. cambridgeassessment.org.uk
Computer-adaptive testing

Testing in which each item, or each block of items, is chosen to suit the student's estimated ability, giving more precision for the same number of items. PISA 2018 used a multistage version for reading (Yamamoto et al., 2018).

  • Yamamoto, K., Shin, H. J., & Khorramdel, L. (2018). Multistage adaptive testing design in international large-scale assessments. Educational Measurement: Issues and Practice, 37(4), 16–27. doi
  • Rasch, G. (1960). Probabilistic models for some intelligence and attainment tests. Danmarks Paedagogiske Institut. archive.org
Consequential validity

The part of the validity argument concerned with the consequences of using scores in a particular way. Messick (1995) treats it as an aspect of construct validity; Kane (2013) holds that negative consequences can make a score use unacceptable.

  • Messick, S. (1995). Validity of psychological assessment: Validation of inferences from persons’ responses and performances as scientific inquiry into score meaning. American Psychologist, 50(9), 741–749. doi
  • Kane, M. T. (2013). Validating the interpretations and uses of test scores. Journal of Educational Measurement, 50(1), 1–73. doi
  • AERA, APA, & NCME. (2014). Standards for educational and psychological testing. American Educational Research Association. testingstandards.net
Construct and construct validity

A construct is "the concept or characteristic that a test is designed to measure" (AERA et al., 2014). Construct validity is the evidence that the scores behave as that concept would predict.

  • Messick, S. (1995). Validity of psychological assessment: Validation of inferences from persons’ responses and performances as scientific inquiry into score meaning. American Psychologist, 50(9), 741–749. doi
  • AERA, APA, & NCME. (2014). Standards for educational and psychological testing. American Educational Research Association. testingstandards.net
Construct underrepresentation and construct-irrelevant variance

Two threats to validity: an assessment too narrow to include important facets of the construct, and one too broad, picking up things that are not the point, such as reading load in science (Messick, 1995).

  • Messick, S. (1995). Validity of psychological assessment: Validation of inferences from persons’ responses and performances as scientific inquiry into score meaning. American Psychologist, 50(9), 741–749. doi
  • Crooks, T. J., Kane, M. T., & Cohen, A. S. (1996). Threats to the valid use of assessments. Assessment in Education: Principles, Policy & Practice, 3(3), 265–286. doi
  • AERA, APA, & NCME. (2014). Standards for educational and psychological testing. American Educational Research Association. testingstandards.net
Content validity

Evidence that the tasks represent the domain being claimed, usually from expert review against a blueprint. In current usage it is one source of evidence, not a separate kind of validity.

  • AERA, APA, & NCME. (2014). Standards for educational and psychological testing. American Educational Research Association. testingstandards.net
  • Messick, S. (1995). Validity of psychological assessment: Validation of inferences from persons’ responses and performances as scientific inquiry into score meaning. American Psychologist, 50(9), 741–749. doi
Cut scores

The marks that separate one reported category, grade or decision from another. They are set by structured expert judgement, and the Standards require the rationale to be documented (AERA et al., 2014).

  • Cizek, G. J., & Bunch, M. B. (2007). Standard setting: A guide to establishing and evaluating performance standards on tests. Sage. sagepub.com
  • AERA, APA, & NCME. (2014). Standards for educational and psychological testing. American Educational Research Association. testingstandards.net
D
Differential item functioning

"For a particular item in a test, a statistical indicator of the extent to which different groups of test takers who are at the same ability level have different frequencies of correct responses" (AERA et al., 2014). A flag for investigation, not proof of bias.

  • AERA, APA, & NCME. (2014). Standards for educational and psychological testing. American Educational Research Association. testingstandards.net
  • Holland, P. W., & Wainer, H. (Eds.). (1993). Differential item functioning. Lawrence Erlbaum Associates. doi
E
E-assessment

Assessment delivered, captured or marked by computer, from on-screen examinations to automatically marked practice. It changes access, security and speed of reporting more than it changes what is being claimed.

  • Yamamoto, K., Shin, H. J., & Khorramdel, L. (2018). Multistage adaptive testing design in international large-scale assessments. Educational Measurement: Issues and Practice, 37(4), 16–27. doi
F
Fairness and bias

Fairness is "the validity of test score interpretations for intended use(s) for individuals from all relevant subgroups" (AERA et al., 2014). Bias is systematic error that makes scores mean different things for different groups.

  • AERA, APA, & NCME. (2014). Standards for educational and psychological testing. American Educational Research Association. testingstandards.net
  • Liang, W., Yuksekgonul, M., Mao, Y., Wu, E., & Zou, J. (2023). GPT detectors are biased against non-native English writers. Patterns, 4(7), 100779. doi
G
Grade inflation

A rise in grades over time without a matching rise in the performance they describe. Strathern (1997) describes the mechanism: once a grade becomes an expectation, it stops discriminating.

  • Strathern, M. (1997). ‘Improving ratings’: Audit in the British University system. European Review, 5(3), 305–321. doi
  • Ofqual. (2020b). Grading GCSEs, AS and A levels in 2021 (Ofqual/20/6720/2). Office of Qualifications and Examinations Regulation. assets.publishing.service.gov.uk
Grading and reporting

Turning performance into symbols for an audience. Brookhart et al. (2016) find that grades combine attainment with effort and behaviour, which makes them useful predictors and poor measurements.

  • Brookhart, S. M., Guskey, T. R., Bowers, A. J., McMillan, J. H., Smith, J. K., Smith, L. F., Stevens, M. T., & Welsh, M. E. (2016). A century of grading research: Meaning and value in the most common educational measure. Review of Educational Research, 86(4), 803–848. doi
  • Kohn, A. (2011). The case against grades. Educational Leadership, 69(3), 28–33. ascd.org
H
High-stakes testing

Testing whose results carry significant consequences for students, teachers or schools, such as certification, admission, funding or public judgement.

  • Campbell, D. T. (1979). Assessing the impact of planned social change. Evaluation and Program Planning, 2(1), 67–90. doi
  • Koretz, D. (2008). Measuring up: What educational testing really tells us. Harvard University Press. jstor.org
  • Au, W. (2007). High-stakes testing and curricular control: A qualitative metasynthesis. Educational Researcher, 36(5), 258–267. doi
I
Inter-rater reliability

The degree to which independent markers reach the same judgement on the same work. Ofqual (2018) reports a related index: the probability that a candidate receives the definitive grade.

  • Ofqual. (2018). Marking consistency metrics: An update (Ofqual/18/6449/2). Office of Qualifications and Examinations Regulation. assets.publishing.service.gov.uk
  • Koretz, D., Stecher, B., Klein, S., & McCaffrey, D. (1994). The Vermont portfolio assessment program: Findings and implications. Educational Measurement: Issues and Practice, 13(3), 5–16. doi
Ipsative referencing

Comparing a student's performance with their own earlier performance rather than with others or a standard. Its home in this series is Formative assessment and feedback (C13).

  • Hughes, G. (2011). Towards a personal best: A case for introducing ipsative assessment in higher education. Studies in Higher Education, 36(3), 353–367. doi
Item banks

A calibrated collection of items "from which a test or test scale's items are selected" (AERA et al., 2014). Banks make multiple forms and adaptive delivery possible, and make exposure and security a standing concern.

  • Rasch, G. (1960). Probabilistic models for some intelligence and attainment tests. Danmarks Paedagogiske Institut. archive.org
  • AERA, APA, & NCME. (2014). Standards for educational and psychological testing. American Educational Research Association. testingstandards.net
  • Yamamoto, K., Shin, H. J., & Khorramdel, L. (2018). Multistage adaptive testing design in international large-scale assessments. Educational Measurement: Issues and Practice, 37(4), 16–27. doi
Item response theory

A model of the relationship between an item, its properties and a person's standing on the construct, placing items and people on one scale, which makes item banking and adaptive testing possible (Rasch, 1960).

  • Rasch, G. (1960). Probabilistic models for some intelligence and attainment tests. Danmarks Paedagogiske Institut. archive.org
  • Yamamoto, K., Shin, H. J., & Khorramdel, L. (2018). Multistage adaptive testing design in international large-scale assessments. Educational Measurement: Issues and Practice, 37(4), 16–27. doi
M
Measurement error

The difference between the score observed and the score a perfect measurement would give. Classical test theory treats an observed score as a true score plus independent random error (AERA et al., 2014).

  • AERA, APA, & NCME. (2014). Standards for educational and psychological testing. American Educational Research Association. testingstandards.net
  • Wiliam, D. (2001). Reliability, validity, and all that jazz. Education 3–13, 29(3), 17–21. discovery.ucl.ac.uk
  • He, Q., Hayes, M., & Wiliam, D. (2011). Classification accuracy in results from Key Stage 2 National Curriculum tests (Ofqual/11/4830). Office of Qualifications and Examinations Regulation. assets.publishing.service.gov.uk
Moderation

Bringing different markers, tasks or cohorts onto a common scale, whether statistically or through meetings in which markers compare judgements against agreed examples. In the Standards it is "a process of relating scores on different tests so that scores have the same relative meaning".

  • AERA, APA, & NCME. (2014). Standards for educational and psychological testing. American Educational Research Association. testingstandards.net
  • Ofqual. (2020b). Grading GCSEs, AS and A levels in 2021 (Ofqual/20/6720/2). Office of Qualifications and Examinations Regulation. assets.publishing.service.gov.uk
N
National examinations

System-wide examinations that certify attainment and usually control progression, run by a public board. In Singapore these are set and administered by the Singapore Examinations and Assessment Board.

  • Ministry of Education, Singapore. (2016, July 13). Changes to PSLE scoring and Secondary One posting [Press release]. moe.gov.sg
  • Singapore Examinations and Assessment Board. (n.d.). Secondary Education Certificate (SEC). Retrieved September 18, 2026, from seab.gov.sg
  • Ministry of Education, Singapore. (n.d.). How will the current PSLE scoring system benefit your child? Retrieved September 18, 2026, from moe.gov.sg
Norm-, criterion- and standards-referencing

Three reference points. Norm-referenced results compare a student with a reference population; criterion-referenced results describe performance against a defined domain (Glaser, 1963); standards-referenced results place students into described levels using cut scores.

  • Glaser, R. (1963). Instructional technology and the measurement of learning outcomes: Some questions. American Psychologist, 18(8), 519–521. doi
  • AERA, APA, & NCME. (2014). Standards for educational and psychological testing. American Educational Research Association. testingstandards.net
P
Performance assessment

Assessment in which students produce or do something extended, such as an experiment, a design or a presentation, and are judged on the performance itself.

  • Wiggins, G. (1990). The case for authentic assessment (ERIC Digest ED328611). ERIC Clearinghouse on Tests, Measurement and Evaluation. files.eric.ed.gov
  • Koretz, D., Stecher, B., Klein, S., & McCaffrey, D. (1994). The Vermont portfolio assessment program: Findings and implications. Educational Measurement: Issues and Practice, 13(3), 5–16. doi
Portfolios

Collections of student work assembled over time, often with reflection. Rich for learning; hard to score consistently across classrooms, as Vermont's programme showed (Koretz et al., 1994).

  • Koretz, D., Stecher, B., Klein, S., & McCaffrey, D. (1994). The Vermont portfolio assessment program: Findings and implications. Educational Measurement: Issues and Practice, 13(3), 5–16. doi
R
Reliability

"The degree to which test scores for a group of test takers are consistent over repeated applications of a measurement procedure" (AERA et al., 2014). It limits, but does not guarantee, validity.

  • AERA, APA, & NCME. (2014). Standards for educational and psychological testing. American Educational Research Association. testingstandards.net
  • Wiliam, D. (2001). Reliability, validity, and all that jazz. Education 3–13, 29(3), 17–21. discovery.ucl.ac.uk
  • Ofqual. (2018). Marking consistency metrics: An update (Ofqual/18/6449/2). Office of Qualifications and Examinations Regulation. assets.publishing.service.gov.uk
S
Score inflation

Rising scores on the test that carries the stakes, without a matching rise on other measures of the same learning (Koretz & Barron, 1998).

  • Koretz, D., & Barron, S. I. (1998). The validity of gains in scores on the Kentucky Instructional Results Information System (KIRIS) (MR-1014-EDU). RAND. rand.org
  • Koretz, D. (2008). Measuring up: What educational testing really tells us. Harvard University Press. jstor.org
  • Koretz, D. (2017). The testing charade: Pretending to make schools better. University of Chicago Press. press.uchicago.edu
Sources of validity evidence

Test content, response processes, internal structure, relations to other variables and the consequences of testing: a menu to draw from according to which assumption in the argument is shakiest (AERA et al., 2014).

  • AERA, APA, & NCME. (2014). Standards for educational and psychological testing. American Educational Research Association. testingstandards.net
Standard error of measurement

"The standard deviation of measurement errors that affect the scores of test takers at a specified test score level" (AERA et al., 2014). It converts reliability into the units of the test, which is what makes it useful at a boundary.

  • AERA, APA, & NCME. (2014). Standards for educational and psychological testing. American Educational Research Association. testingstandards.net
Standard setting

"The process, often judgment based, of setting cut scores using a structured procedure" that maps scores onto described performance levels (AERA et al., 2014). Cizek and Bunch (2007) describe the main methods.

  • Cizek, G. J., & Bunch, M. B. (2007). Standard setting: A guide to establishing and evaluating performance standards on tests. Sage. sagepub.com
  • AERA, APA, & NCME. (2014). Standards for educational and psychological testing. American Educational Research Association. testingstandards.net
T
Table of specifications

A blueprint showing how many items or marks each topic and each kind of thinking receives. Reading one tells you what the paper can and cannot support a claim about.

  • AERA, APA, & NCME. (2014). Standards for educational and psychological testing. American Educational Research Association. testingstandards.net
Teaching to the test

Directing teaching towards the sample of the domain that the test happens to contain. It raises scores; whether it raises learning depends on how well the test samples the domain (Koretz, 2008).

  • Koretz, D. (2008). Measuring up: What educational testing really tells us. Harvard University Press. jstor.org
  • Au, W. (2007). High-stakes testing and curricular control: A qualitative metasynthesis. Educational Researcher, 36(5), 258–267. doi
  • Koretz, D. (2017). The testing charade: Pretending to make schools better. University of Chicago Press. press.uchicago.edu
Two-lane approach

Splitting a programme's assessment into a secured, in-person lane that carries the decisions and an open lane where learning, feedback and supported use of AI happen (Bridgeman et al., 2024).

  • Bridgeman, A., Weeks, R., & Liu, D. (2024, November 26). Aligning our assessments to the age of generative AI. Teaching@Sydney, University of Sydney. educational-innovation.sydney.edu.au
  • Lodge, J. M., Howard, S., Bearman, M., Dawson, P., & Associates. (2023). Assessment reform for the age of artificial intelligence. Tertiary Education Quality and Standards Agency. teqsa.gov.au
V
Validity

"The degree to which evidence and theory support the interpretations of test scores for proposed uses of tests" (AERA et al., 2014). It belongs to an interpretation for a use, never to the instrument alone.

  • Messick, S. (1995). Validity of psychological assessment: Validation of inferences from persons’ responses and performances as scientific inquiry into score meaning. American Psychologist, 50(9), 741–749. doi
  • Kane, M. T. (2013). Validating the interpretations and uses of test scores. Journal of Educational Measurement, 50(1), 1–73. doi
  • AERA, APA, & NCME. (2014). Standards for educational and psychological testing. American Educational Research Association. testingstandards.net
Validity argument

The set of claims, inferences and supporting evidence that links a performance to a decision. Kane (2013) calls the statement of what is claimed the interpretation/use argument, and validation its evaluation.

  • Kane, M. T. (2013). Validating the interpretations and uses of test scores. Journal of Educational Measurement, 50(1), 1–73. doi
  • Cook, D. A., Brydges, R., Ginsburg, S., & Hatala, R. (2015). A contemporary approach to validity arguments: A practical guide to Kane’s framework. Medical Education, 49(6), 560–575. doi
  • Crooks, T. J., Kane, M. T., & Cohen, A. S. (1996). Threats to the valid use of assessments. Assessment in Education: Principles, Policy & Practice, 3(3), 265–286. doi
W
Washback

The effect a test has on teaching and learning before it is taken. Alderson and Wall (1993) argued it is neither automatic nor simple, and its size varies with the stakes and the teacher (Alderson & Hamp-Lyons, 1996).

  • Alderson, J. C., & Wall, D. (1993). Does washback exist? Applied Linguistics, 14(2), 115–129. doi
  • Alderson, J. C., & Hamp-Lyons, L. (1996). TOEFL preparation courses: A study of washback. Language Testing, 13(3), 280–297. doi
Singapore