Gerald Ajam

Library / Glossary / The craft glossary

Evaluative research glossary

48 terms from R05, Evaluative research: usability testing, surveys and UX metrics — 24 defined in the guide itself and 24 more from the field around it. Every term the guide teaches links to the slide that teaches it.

The whole craft glossary

R05 · What to build

Evaluative research: usability testing, surveys and UX metrics

Does what we made work for the people using it?

Every term below is defined in the words of evaluative research, guide R05 of craft guides for educational technologists, and opens the guide at the slide where it is taught. 24 of the 48 are the field’s vocabulary rather than the guide’s own: words a reader will meet around this subject, defined here because the guide assumes them. 4 terms are also defined by another guide in the series; where the two differ, both wordings are given. The whole craft glossary holds all of them together.

TermDefinitionReferred to inRead further
A
Acquiescence bias

The tendency of some respondents to agree with a statement whatever it says (Pew Research Center, 2021).

  • Schuman, H., & Presser, S. (1981). Questions and answers in attitude surveys: Experiments on question form, wording, and context. Academic Press. archive.org
  • Krosnick, J. A. (1991). Response strategies for coping with the cognitive demands of attitude measures in surveys. Applied Cognitive Psychology, 5(3), 213–236. doi
  • Pew Research Center. (2021). Writing survey questions. pewresearch.org
Adjusted-Wald interval

A confidence interval for a proportion that behaves well with small samples (Sauro & Lewis, 2016).

  • Sauro, J., & Lewis, J. R. (2016). Quantifying the user experience: Practical statistics for user research (2nd ed.). Morgan Kaufmann. shop.elsevier.com
  • Agresti, A., & Coull, B. A. (1998). Approximate is better than "exact" for interval estimation of binomial proportions. The American Statistician, 52(2), 119–126. doi
Assent

A child's own agreement to take part, sought alongside consent from a responsible adult (British Educational Research Association, 2024).

R04 A child's own agreement to take part in research, sought alongside a guardian's consent (British Educational Research Association, 2024).

R13 Agreement to take part given by children too young to consent; principles of consent apply to children and young people as well as to adults (British Educational Research Association, 2024).

  • British Educational Research Association. (2024). Ethical guidelines for educational research (5th ed.). bera.ac.uk
  • Lundy, L. (2007). ‘Voice’ is not enough: Conceptualising Article 12 of the United Nations Convention on the Rights of the Child. British Educational Research Journal, 33(6), 927–942. doi
  • United Nations. (1989). Convention on the Rights of the Child (General Assembly resolution 44/25). ohchr.org
Attitudinal method

A method that records what people say: their views, ratings and feelings.

  • Rohrer, C. (2022). When to use which user-experience research methods. Nielsen Norman Group. nngroup.com
B
Behavioural method

A method that records what people do: their actions, successes and errors.

  • Rohrer, C. (2022). When to use which user-experience research methods. Nielsen Norman Group. nngroup.com
C
Card sorting

A method in which participants group labelled cards in a way that makes sense to them, used to learn how users expect content to be organised and named.

  • R05Evaluative researchFrom the field · not in the guide
  • Rohrer, C. (2022). When to use which user-experience research methods. Nielsen Norman Group. nngroup.com
  • Spencer, D. (2009). Card sorting: Designing usable categories. Rosenfeld Media.
Closed question

A survey question answered by choosing from a fixed set of options. It is easy to count, but it can only return the answers its writer thought of.

  • R05Evaluative researchFrom the field · not in the guide
  • Pew Research Center. (2021). Writing survey questions. pewresearch.org
  • Schuman, H., & Presser, S. (1981). Questions and answers in attitude surveys: Experiments on question form, wording, and context. Academic Press. archive.org
Cognitive interview

A survey pretesting method in which a few respondents answer draft questions while explaining how they understood them and how they arrived at each answer.

  • R05Evaluative researchFrom the field · not in the guide
  • Willis, G. B. (2005). Cognitive interviewing: A tool for improving questionnaire design. SAGE. doi
Cognitive walkthrough

An inspection method in which reviewers step through a task as a first-time user would, asking at each step whether that user would know what to do and see that it had worked.

  • R05Evaluative researchFrom the field · not in the guide
  • Wharton, C., Rieman, J., Lewis, C., & Polson, P. (1994). The cognitive walkthrough method: A practitioner's guide. In J. Nielsen & R. L. Mack (Eds.), Usability inspection methods (pp. 105–140). Wiley.
Context of use

The users, goals, tasks, equipment and environment in which a product is used (International Organization for Standardization, 2018).

  • International Organization for Standardization. (2018). Ergonomics of human-system interaction — Part 11: Usability: Definitions and concepts (ISO Standard No. 9241-11:2018). iso.org
D
Double-barrelled question

A question that asks about two things at once, such as whether a tool is "quick and easy", so an answer cannot be tied to either.

  • R05Evaluative researchFrom the field · not in the guide
E
Error rate

The number of mistakes users make while attempting a task, reported per task or per opportunity for error.

  • R05Evaluative researchFrom the field · not in the guide
  • Sauro, J., & Lewis, J. R. (2016). Quantifying the user experience: Practical statistics for user research (2nd ed.). Morgan Kaufmann. shop.elsevier.com
Evaluator effect

The finding that different evaluators using the same method on the same system report different problems (Hertzum & Jacobsen, 2001).

  • Hertzum, M., & Jacobsen, N. E. (2001). The evaluator effect: A chilling fact about usability evaluation methods. International Journal of Human–Computer Interaction, 13(4), 421–443. doi
Eye tracking

Recording where on a screen a person looks, for how long and in what order, using a device that follows the movement of their eyes.

  • R05Evaluative researchFrom the field · not in the guide
  • Rohrer, C. (2022). When to use which user-experience research methods. Nielsen Norman Group. nngroup.com
  • Nielsen, J., & Pernice, K. (2010). Eyetracking web usability. New Riders.
F
Five-user rule

The advice to test with five users, from a model in which one user reveals about 31% of problems, so five find about 85% on average, though any five may not (Nielsen, 2000).

  • Nielsen, J., & Landauer, T. K. (1993). A mathematical model of the finding of usability problems. In Proceedings of the INTERACT '93 and CHI '93 Conference on Human Factors in Computing Systems (pp. 206–213). ACM. doi
  • Nielsen, J. (2000). Why you only need to test with 5 users. Nielsen Norman Group. nngroup.com
  • Faulkner, L. (2003). Beyond the five-user assumption: Benefits of increased sample sizes in usability testing. Behavior Research Methods, Instruments, & Computers, 35(3), 379–383. doi
Formative and summative testing

Formative testing looks for problems to fix while a design is still changing. Summative testing measures how well a finished design performs, often against a target or an earlier version.

  • R05Evaluative researchFrom the field · not in the guide
  • Sauro, J., & Lewis, J. R. (2016). Quantifying the user experience: Practical statistics for user research (2nd ed.). Morgan Kaufmann. shop.elsevier.com
  • Rubin, J., & Chisnell, D. (2008). Handbook of usability testing: How to plan, design, and conduct effective tests (2nd ed.). Wiley.
G
Goals, signals, metrics

A simple process for choosing product metrics: name the goals, find the signals that show success or failure, and turn the signals into metrics (Rodden et al., 2010).

  • Rodden, K., Hutchinson, H., & Fu, X. (2010). Measuring the user experience on a large scale: User-centered metrics for web applications. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems (pp. 2395–2398). ACM. doi
Guerrilla testing

Quick, informal usability testing with whoever is to hand, such as colleagues or passers-by. It trades a careful sample for speed and low cost.

  • R05Evaluative researchFrom the field · not in the guide
  • Krug, S. (2010). Rocket surgery made easy: The do-it-yourself guide to finding and fixing usability problems. New Riders.
H
HEART

Happiness, engagement, adoption, retention and task success: categories for product metrics (Rodden et al., 2010).

  • Rodden, K., Hutchinson, H., & Fu, X. (2010). Measuring the user experience on a large scale: User-centered metrics for web applications. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems (pp. 2395–2398). ACM. doi
I
Intercept survey

A short survey shown to people while they are using a website or product, so that answers are gathered in the moment of use.

  • R05Evaluative researchFrom the field · not in the guide
  • Rohrer, C. (2022). When to use which user-experience research methods. Nielsen Norman Group. nngroup.com
L
Lab and field testing

The choice between testing in a lab and testing where the product is used; the evidence is mixed, and the setting has to be chosen on purpose (Kjeldskov & Skov, 2014).

  • Kjeldskov, J., Skov, M. B., Als, B. S., & Høegh, R. T. (2004). Is it worth the hassle? Exploring the added value of evaluating the usability of context-aware mobile systems in the field. In S. Brewster & M. Dunlop (Eds.), Mobile human-computer interaction: MobileHCI 2004 (pp. 61–73). Springer. doi
  • Kjeldskov, J., & Skov, M. B. (2014). Was it worth the hassle? Ten years of mobile HCI research discussions on lab and field evaluations. In Proceedings of the 16th International Conference on Human-Computer Interaction with Mobile Devices and Services (pp. 43–52). ACM. doi
Leading question

A question whose wording suggests the answer the asker hopes for.

  • Schuman, H., & Presser, S. (1981). Questions and answers in attitude surveys: Experiments on question form, wording, and context. Academic Press. archive.org
  • Pew Research Center. (2021). Writing survey questions. pewresearch.org
Learnability

How quickly and easily a new user can reach competent use of a product. It is measured by how performance improves over repeated attempts.

  • R05Evaluative researchFrom the field · not in the guide
  • Nielsen, J. (1993). Usability engineering. Academic Press. doi
Likert scale

A set of statements, each rated on an ordered scale from strongly disagree to strongly agree, with the ratings combined into one score. The name is often used loosely for a single such item.

  • R05Evaluative researchFrom the field · not in the guide
  • Likert, R. (1932). A technique for the measurement of attitudes. Archives of Psychology, 22(140), 1–55.
M
Moderated and unmoderated testing

In a moderated test a researcher guides each participant through the session, in person or remotely. In an unmoderated test participants work through the tasks alone, usually through an online tool.

  • R05Evaluative researchFrom the field · not in the guide
  • Rohrer, C. (2022). When to use which user-experience research methods. Nielsen Norman Group. nngroup.com
  • Rubin, J., & Chisnell, D. (2008). Handbook of usability testing: How to plan, design, and conduct effective tests (2nd ed.). Wiley.
N
Net Promoter Score

The share of people rating their likelihood to recommend at nine or ten, minus the share rating it six or below (Reichheld, 2003).

  • Reichheld, F. F. (2003). The one number you need to grow. Harvard Business Review, 81(12), 46–54. hbr.org
  • Keiningham, T. L., Cooil, B., Andreassen, T. W., & Aksoy, L. (2007). A longitudinal examination of Net Promoter and firm revenue growth. Journal of Marketing, 71(3), 39–51. doi
Nonresponse bias

Error that arises when the people who answer a survey differ from those who do not.

  • Groves, R. M., & Peytcheva, E. (2008). The impact of nonresponse rates on nonresponse bias: A meta-analysis. Public Opinion Quarterly, 72(2), 167–189. doi
  • Davern, M. (2013). Nonresponse rates are a problematic indicator of nonresponse bias in survey research. Health Services Research, 48(3), 905–912. doi
P
Pilot test

A trial run of a study with one or two people, held to find faults in the tasks, questions, timing and equipment before the real sessions begin.

  • R05Evaluative researchFrom the field · not in the guide
  • Rubin, J., & Chisnell, D. (2008). Handbook of usability testing: How to plan, design, and conduct effective tests (2nd ed.). Wiley.
Q
Question order effect

A change in how people answer a survey question caused by the questions that came before it.

  • R05Evaluative researchFrom the field · not in the guide
  • Schuman, H., & Presser, S. (1981). Questions and answers in attitude surveys: Experiments on question form, wording, and context. Academic Press. archive.org
  • Pew Research Center. (2021). Writing survey questions. pewresearch.org
R
Response order effect

A change in which answer option people choose caused by the order in which the options are presented, such as favouring the first option in a written list.

  • R05Evaluative researchFrom the field · not in the guide
  • Krosnick, J. A. (1991). Response strategies for coping with the cognitive demands of attitude measures in surveys. Applied Cognitive Psychology, 5(3), 213–236. doi
  • Pew Research Center. (2021). Writing survey questions. pewresearch.org
Response rate

The share of the people invited to take a survey who complete it. A low rate is a warning of possible nonresponse bias, not proof of it.

  • R05Evaluative researchFrom the field · not in the guide
  • Groves, R. M., & Peytcheva, E. (2008). The impact of nonresponse rates on nonresponse bias: A meta-analysis. Public Opinion Quarterly, 72(2), 167–189. doi
  • Davern, M. (2013). Nonresponse rates are a problematic indicator of nonresponse bias in survey research. Health Services Research, 48(3), 905–912. doi
S
Sampling frame

The list from which a survey sample is drawn, such as a register of teachers. People missing from the list cannot be selected, whatever the sample size.

  • R05Evaluative researchFrom the field · not in the guide
  • Groves, R. M., Fowler, F. J., Jr., Couper, M. P., Lepkowski, J. M., Singer, E., & Tourangeau, R. (2009). Survey methodology (2nd ed.). Wiley.
Satisficing

Giving an answer that is good enough rather than accurate, to save effort (Krosnick, 1991).

  • Krosnick, J. A. (1991). Response strategies for coping with the cognitive demands of attitude measures in surveys. Applied Cognitive Psychology, 5(3), 213–236. doi
  • Simon, H. A. (1956). Rational choice and the structure of the environment. Psychological Review, 63(2), 129–138. doi
Severity rating

A grade given to each usability problem, usually based on how many users it affects and how badly, used to decide which problems to fix first.

  • R05Evaluative researchFrom the field · not in the guide
  • Sauro, J., & Lewis, J. R. (2016). Quantifying the user experience: Practical statistics for user research (2nd ed.). Morgan Kaufmann. shop.elsevier.com
  • Dumas, J. S., & Redish, J. C. (1999). A practical guide to usability testing (Rev. ed.). Intellect.
Single Ease Question (SEQ)

A one-item questionnaire given straight after a task, asking the user to rate how difficult or easy it was on a seven-point scale.

  • R05Evaluative researchFrom the field · not in the guide
  • Sauro, J., & Lewis, J. R. (2016). Quantifying the user experience: Practical statistics for user research (2nd ed.). Morgan Kaufmann. shop.elsevier.com
  • Sauro, J., & Dumas, J. S. (2009). Comparison of three one-question, post-task usability questionnaires. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems (pp. 1599–1608). ACM. doi
Standardised questionnaire

A questionnaire with fixed wording, response scale and scoring, whose reliability and validity have been tested, so that scores can be compared across studies and products.

  • R05Evaluative researchFrom the field · not in the guide
  • Sauro, J., & Lewis, J. R. (2016). Quantifying the user experience: Practical statistics for user research (2nd ed.). Morgan Kaufmann. shop.elsevier.com
Survey fatigue

Falling willingness to respond as people receive more surveys (Porter et al., 2004).

  • Porter, S. R., Whitcomb, M. E., & Weitzer, W. H. (2004). Multiple surveys of students and survey fatigue. New Directions for Institutional Research, 2004(121), 63–73. doi
SUS

The System Usability Scale: ten items giving one score from 0 to 100 (Brooke, 1996).

  • Brooke, J. (1996). SUS: A 'quick and dirty' usability scale. In P. W. Jordan, B. Thomas, B. A. Weerdmeester, & I. L. McClelland (Eds.), Usability evaluation in industry (pp. 189–194). Taylor & Francis. doi
  • Bangor, A., Kortum, P. T., & Miller, J. T. (2008). An empirical evaluation of the System Usability Scale. International Journal of Human–Computer Interaction, 24(6), 574–594. doi
  • Bangor, A., Kortum, P., & Miller, J. (2009). Determining what individual SUS scores mean: Adding an adjective rating scale. Journal of Usability Studies, 4(3), 114–123. dl.acm.org
T
Task scenario

A short, realistic situation given to a usability test participant that states a goal to reach, without naming the steps or the words used on screen.

  • R05Evaluative researchFrom the field · not in the guide
  • Rubin, J., & Chisnell, D. (2008). Handbook of usability testing: How to plan, design, and conduct effective tests (2nd ed.). Wiley.
Task success

Whether a participant completes a test task, the usual measure of effectiveness.

  • International Organization for Standardization. (2018). Ergonomics of human-system interaction — Part 11: Usability: Definitions and concepts (ISO Standard No. 9241-11:2018). iso.org
  • Sauro, J., & Lewis, J. R. (2016). Quantifying the user experience: Practical statistics for user research (2nd ed.). Morgan Kaufmann. shop.elsevier.com
  • Sauro, J. (2011). What is a good task-completion rate? MeasuringU. measuringu.com
Test plan

The document agreed before a usability study that sets out its purpose, research questions, participants, tasks, measures and schedule.

  • R05Evaluative researchFrom the field · not in the guide
  • Rubin, J., & Chisnell, D. (2008). Handbook of usability testing: How to plan, design, and conduct effective tests (2nd ed.). Wiley.
Thinking aloud

Asking participants to say what they are thinking while they work on a task (Nielsen, 2012).

  • Ericsson, K. A., & Simon, H. A. (1993). Protocol analysis: Verbal reports as data (Rev. ed.). MIT Press. doi
  • Fox, M. C., Ericsson, K. A., & Best, R. (2011). Do procedures for verbal reporting of thinking have to be reactive? A meta-analysis and recommendations for best reporting methods. Psychological Bulletin, 137(2), 316–344. doi
  • Nielsen, J. (2012). Thinking aloud: The #1 usability tool. Nielsen Norman Group. nngroup.com
Time on task

How long a participant takes to complete a task, a measure of efficiency.

R15 Time a student is estimated to have spent on an activity, usually inferred from gaps between logged actions (Kovanović et al., 2015).

  • International Organization for Standardization. (2018). Ergonomics of human-system interaction — Part 11: Usability: Definitions and concepts (ISO Standard No. 9241-11:2018). iso.org
  • Kovanović, V., Gašević, D., Dawson, S., Joksimović, S., Baker, R. S., & Hatala, M. (2015). Penetrating the black box of time-on-task estimation. In Proceedings of the Fifth International Conference on Learning Analytics and Knowledge (pp. 184–193). ACM. doi
  • Sauro, J., & Lewis, J. R. (2016). Quantifying the user experience: Practical statistics for user research (2nd ed.). Morgan Kaufmann. shop.elsevier.com
Tree testing

A test of a menu structure in which participants are shown only the text hierarchy and asked where they would look for given items.

  • R05Evaluative researchFrom the field · not in the guide
  • Rohrer, C. (2022). When to use which user-experience research methods. Nielsen Norman Group. nngroup.com
U
UMUX-LITE

A two-item usability questionnaire that tracks SUS closely (Lewis et al., 2013).

  • Lewis, J. R., Utesch, B. S., & Maher, D. E. (2013). UMUX-LITE: When there's no time for the SUS. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems (pp. 2099–2102). ACM. doi
  • Finstad, K. (2010). The usability metric for user experience. Interacting with Computers, 22(5), 323–327. doi
Usability

The extent to which a product can be used by specified users to achieve specified goals with effectiveness, efficiency and satisfaction in a specified context of use (International Organization for Standardization, 2018).

  • International Organization for Standardization. (2018). Ergonomics of human-system interaction — Part 11: Usability: Definitions and concepts (ISO Standard No. 9241-11:2018). iso.org
Usability test

A test that samples people, tasks and conditions: it shows how some people did some tasks in one setting, and holds for what was specified and no further.

  • International Organization for Standardization. (2018). Ergonomics of human-system interaction — Part 11: Usability: Definitions and concepts (ISO Standard No. 9241-11:2018). iso.org
  • Rohrer, C. (2022). When to use which user-experience research methods. Nielsen Norman Group. nngroup.com
W
Web Content Accessibility Guidelines (WCAG)

The Web Content Accessibility Guidelines; version 2.2 became a W3C Recommendation in October 2023, and meeting it does not address every user need (World Wide Web Consortium, 2023).

R08 The World Wide Web Consortium's accessibility standard; version 2.2 sets minimums such as 4.5:1 contrast for normal text and 24 by 24 CSS pixel targets (World Wide Web Consortium, 2023).

  • World Wide Web Consortium. (2023). Web Content Accessibility Guidelines (WCAG) 2.2 (W3C Recommendation). w3.org
  • W3C Web Accessibility Initiative. (2024). Involving users in evaluating web accessibility. w3.org
Singapore