Library / Glossary / The context glossary
Assessment systems glossary
42 terms from C14, Assessment systems: validity, reliability and high stakes — each defined in the primer's own words. Every term the primer teaches links to the slide that teaches it.
C14 · Curriculum and assessment
Assessment systems: validity, reliability and high stakes
“A score measures what the item says it measures.”
Every term below is defined in the words of assessment systems, primer C14 of understanding the context of education, and opens the primer at the slide where it is taught. 1 term is also defined by another primer in the series; where the two differ, both wordings are given. The whole context glossary holds all of them together.
| Term | Definition | Referred to in | Read further |
|---|---|---|---|
| A | |||
| Academic integrity | The expectation that work submitted for assessment is the student's own and that sources and tools are acknowledged. It is a property of the whole assessment design, not only of a rule in a handbook. |
| |
| Achievement Levels | Singapore's primary examination scoring from the 2021 cohort: each subject is reported in one of eight Achievement Levels, and the four subjects sum to a score between 4 and 32 (Ministry of Education, Singapore, 2016). |
| |
| AI and assessment design | Redesigning what is assessed, and where, so that the inference from work to student still holds when a capable tool is available. Lodge et al. (2023) argue for authentic engagement with AI plus security at a few meaningful points. |
| |
| Assessment literacy | Knowing what assessments can and cannot tell you, and how to build and read them. The term was popularised by Stiggins (1991), who argued that decision-makers need it as much as teachers do. |
| |
| Audit test | A second measure of the same learning, taken by the same students with no stakes attached; the standard way to detect score inflation, and rarely done. |
| |
| Authentic assessment | Assessment that "directly examines student performance on worthy intellectual tasks" rather than using indirect proxy items (Wiggins, 1990). |
| |
| C | |||
| Campbell's law | The more any quantitative social indicator is used for social decision-making, the more subject it will be to corruption pressures and the more apt it will be to distort the processes it monitors (Campbell, 1979). C01 When test scores become the goal of the teaching process, they both lose their value as indicators of educational status and distort the educational process in undesirable ways (Campbell, 1979). C17 The more a quantitative social indicator is used for social decision-making, the more subject it is to corruption pressures and the more apt to distort the social processes it is intended to monitor (Campbell, 1979). | ||
| Classical test theory | A model that treats an observed score as the sum of a true score and an independent random error component, giving one error figure for the whole test (AERA et al., 2014). |
| |
| Comparative judgement | A method in which assessors repeatedly choose the better of two pieces of work, and a scale is built from the pattern of choices rather than from a rubric (Pollitt, 2012). |
| |
| Computer-adaptive testing | Testing in which each item, or each block of items, is chosen to suit the student's estimated ability, giving more precision for the same number of items. PISA 2018 used a multistage version for reading (Yamamoto et al., 2018). |
| |
| Consequential validity | The part of the validity argument concerned with the consequences of using scores in a particular way. Messick (1995) treats it as an aspect of construct validity; Kane (2013) holds that negative consequences can make a score use unacceptable. |
| |
| Construct and construct validity | A construct is "the concept or characteristic that a test is designed to measure" (AERA et al., 2014). Construct validity is the evidence that the scores behave as that concept would predict. |
| |
| Construct underrepresentation and construct-irrelevant variance | Two threats to validity: an assessment too narrow to include important facets of the construct, and one too broad, picking up things that are not the point, such as reading load in science (Messick, 1995). |
| |
| Content validity | Evidence that the tasks represent the domain being claimed, usually from expert review against a blueprint. In current usage it is one source of evidence, not a separate kind of validity. |
| |
| Cut scores | The marks that separate one reported category, grade or decision from another. They are set by structured expert judgement, and the Standards require the rationale to be documented (AERA et al., 2014). |
| |
| D | |||
| Differential item functioning | "For a particular item in a test, a statistical indicator of the extent to which different groups of test takers who are at the same ability level have different frequencies of correct responses" (AERA et al., 2014). A flag for investigation, not proof of bias. |
| |
| E | |||
| E-assessment | Assessment delivered, captured or marked by computer, from on-screen examinations to automatically marked practice. It changes access, security and speed of reporting more than it changes what is being claimed. |
| |
| F | |||
| Fairness and bias | Fairness is "the validity of test score interpretations for intended use(s) for individuals from all relevant subgroups" (AERA et al., 2014). Bias is systematic error that makes scores mean different things for different groups. |
| |
| G | |||
| Grade inflation | A rise in grades over time without a matching rise in the performance they describe. Strathern (1997) describes the mechanism: once a grade becomes an expectation, it stops discriminating. |
| |
| Grading and reporting | Turning performance into symbols for an audience. Brookhart et al. (2016) find that grades combine attainment with effort and behaviour, which makes them useful predictors and poor measurements. |
| |
| H | |||
| High-stakes testing | Testing whose results carry significant consequences for students, teachers or schools, such as certification, admission, funding or public judgement. |
| |
| I | |||
| Inter-rater reliability | The degree to which independent markers reach the same judgement on the same work. Ofqual (2018) reports a related index: the probability that a candidate receives the definitive grade. |
| |
| Ipsative referencing | Comparing a student's performance with their own earlier performance rather than with others or a standard. Its home in this series is Formative assessment and feedback (C13). |
| |
| Item banks | A calibrated collection of items "from which a test or test scale's items are selected" (AERA et al., 2014). Banks make multiple forms and adaptive delivery possible, and make exposure and security a standing concern. |
| |
| Item response theory | A model of the relationship between an item, its properties and a person's standing on the construct, placing items and people on one scale, which makes item banking and adaptive testing possible (Rasch, 1960). |
| |
| M | |||
| Measurement error | The difference between the score observed and the score a perfect measurement would give. Classical test theory treats an observed score as a true score plus independent random error (AERA et al., 2014). |
| |
| Moderation | Bringing different markers, tasks or cohorts onto a common scale, whether statistically or through meetings in which markers compare judgements against agreed examples. In the Standards it is "a process of relating scores on different tests so that scores have the same relative meaning". |
| |
| N | |||
| National examinations | System-wide examinations that certify attainment and usually control progression, run by a public board. In Singapore these are set and administered by the Singapore Examinations and Assessment Board. |
| |
| Norm-, criterion- and standards-referencing | Three reference points. Norm-referenced results compare a student with a reference population; criterion-referenced results describe performance against a defined domain (Glaser, 1963); standards-referenced results place students into described levels using cut scores. |
| |
| P | |||
| Performance assessment | Assessment in which students produce or do something extended, such as an experiment, a design or a presentation, and are judged on the performance itself. |
| |
| Portfolios | Collections of student work assembled over time, often with reflection. Rich for learning; hard to score consistently across classrooms, as Vermont's programme showed (Koretz et al., 1994). |
| |
| R | |||
| Reliability | "The degree to which test scores for a group of test takers are consistent over repeated applications of a measurement procedure" (AERA et al., 2014). It limits, but does not guarantee, validity. |
| |
| S | |||
| Score inflation | Rising scores on the test that carries the stakes, without a matching rise on other measures of the same learning (Koretz & Barron, 1998). |
| |
| Sources of validity evidence | Test content, response processes, internal structure, relations to other variables and the consequences of testing: a menu to draw from according to which assumption in the argument is shakiest (AERA et al., 2014). |
| |
| Standard error of measurement | "The standard deviation of measurement errors that affect the scores of test takers at a specified test score level" (AERA et al., 2014). It converts reliability into the units of the test, which is what makes it useful at a boundary. |
| |
| Standard setting | "The process, often judgment based, of setting cut scores using a structured procedure" that maps scores onto described performance levels (AERA et al., 2014). Cizek and Bunch (2007) describe the main methods. |
| |
| T | |||
| Table of specifications | A blueprint showing how many items or marks each topic and each kind of thinking receives. Reading one tells you what the paper can and cannot support a claim about. |
| |
| Teaching to the test | Directing teaching towards the sample of the domain that the test happens to contain. It raises scores; whether it raises learning depends on how well the test samples the domain (Koretz, 2008). |
| |
| Two-lane approach | Splitting a programme's assessment into a secured, in-person lane that carries the decisions and an open lane where learning, feedback and supported use of AI happen (Bridgeman et al., 2024). |
| |
| V | |||
| Validity | "The degree to which evidence and theory support the interpretations of test scores for proposed uses of tests" (AERA et al., 2014). It belongs to an interpretation for a use, never to the instrument alone. |
| |
| Validity argument | The set of claims, inferences and supporting evidence that links a performance to a decision. Kane (2013) calls the statement of what is claimed the interpretation/use argument, and validation its evaluation. |
| |
| W | |||
| Washback | The effect a test has on teaching and learning before it is taken. Alderson and Wall (1993) argued it is neither automatic nor simple, and its size varies with the stakes and the teacher (Alderson & Hamp-Lyons, 1996). | ||
No term matches. Try fewer letters.
Nearby