What Does 'Normed Assessment' Mean?
A normed assessment is one in which an individual's score is interpreted relative to a reference population, known as the norm group. When a personality assessment reports that someone scores at the 70th percentile on conscientiousness, that statement is only meaningful because the score has been compared against the distribution of conscientiousness scores in a defined reference population. Norming is the process of collecting that reference data and establishing the conversion rules that translate raw scores into interpretable, population-relative figures. Without a norm group, a raw score carries little interpretive meaning about where a person stands relative to others.
Why Norming Matters
The purpose of norming an assessment is to make scores interpretable. A raw score of 42 on a personality scale tells a practitioner nothing on its own. A normed score that places that 42 at the 68th percentile of a global working adult population tells the practitioner that this individual scored at or above roughly two-thirds of the relevant comparison group on that dimension. That is actionable information for selection, development, and coaching decisions.
Kline (1994), a widely used reference text on psychological testing, treats adequate norming as a prerequisite for professional use of any psychological instrument. The authority on this point is the Standards for Educational and Psychological Testing, which specify that technical documentation for a professionally used assessment must describe the norm group in sufficient detail for users to evaluate whether the norms are appropriate for their intended application. Norm group adequacy is one of the most important and most commonly neglected quality criteria for assessment instruments in applied settings.
Norm-Referenced vs Criterion-Referenced Interpretation
Norm-referenced and criterion-referenced are two fundamentally different interpretive frameworks, and understanding the distinction matters for practitioners choosing between assessment tools.
A norm-referenced interpretation answers the question: how does this person compare to others? The score is expressed relative to the distribution of scores in the norm group, typically as a percentile, a standard score, or a sten score. Norm-referenced interpretation is appropriate when the purpose is to differentiate among candidates or identify where an individual stands on a dimension relative to a relevant population. Virtually all predictive validity research on personality and cognitive ability assessments assumes norm-referenced scoring.
A criterion-referenced interpretation answers the question: does this person meet a defined standard? The score is compared against a fixed threshold rather than a population distribution. Driving tests and professional licensing examinations are criterion-referenced. The question is whether a specific competency level has been reached, not how the candidate compares to other test-takers. Criterion-referenced interpretation is appropriate when there is a meaningful, externally validated minimum standard that must be met regardless of how others perform.
Most talent assessments used in selection and development are norm-referenced. A criterion-referenced interpretation of a personality score would require establishing what score level constitutes adequate conscientiousness for a specific role, a standard that is difficult to set defensibly without extensive local validation data.
What Makes a Good Norm Group?
The interpretive value of a normed assessment depends entirely on the quality and relevance of the norm group. Three properties determine whether a norm group is adequate for a specific application.
Size. A norm group must be large enough to produce stable score distributions across the full range of the scale, and large enough to support stable estimates within relevant subgroups. Small norm groups produce unreliable percentile estimates, particularly at the extremes of the distribution where decisions are often most consequential. Professional standards treat several hundred participants as a minimum floor for a single overall norm, but robust occupational instruments typically draw on samples in the thousands or tens of thousands so that percentile estimates remain stable at the tails and within demographic and occupational subgroups.
Representativeness. The norm group must represent the population against which scores will be interpreted. A personality assessment normed on university undergraduates will produce percentile scores that are systematically misleading when applied to senior executives or manufacturing workers, because the score distribution in those populations is different. The Standards require that technical documentation specify the demographic characteristics, such as age, gender, education, occupation, nationality, of the norm group so that users can evaluate its appropriateness.
Currency. Norm groups can become outdated as population-level characteristics shift over time. Assessments that have not been re-normed in decades may produce percentile interpretations that no longer accurately reflect the current distribution of scores in the relevant population. Responsible publishers update their norms periodically, and technical manuals should document when the normative data were collected.
Norming, Fairness, and Defensible Decisions
Norm choice is not only a technical matter, it is a fairness and compliance matter, and for HR buyers in the United States it bears directly on legal defensibility. Score distributions on personality and ability assessments can differ across demographic subgroups, and the norm against which candidates are compared shapes how those differences translate into selection outcomes. An ill-fitting or outdated norm can introduce or amplify subgroup differences that surface as adverse impact.
Under the Uniform Guidelines on Employee Selection Procedures (EEOC, 1978), selection tools that produce adverse impact against a protected group must be shown to be job-related and consistent with business necessity. The Standards (AERA, APA, & NCME, 2014) and the SIOP Principles for the Validation and Use of Personnel Selection Procedures (2018) both treat appropriate norming and evidence of measurement equivalence across groups as part of the validity and fairness evidence a defensible program should hold. A relevant, current, and representative norm group is therefore not a nicety. It is part of what makes a selection decision fair to candidates and defensible to a regulator.
This also has a practical implication for how scores are used. Norm-referenced percentiles are designed to inform decisions, not to serve as blind automatic cutoffs. Setting a hard percentile gate without local validation can both weaken the job-relatedness of the decision and increase adverse-impact risk. Percentiles are most defensible when they are combined with role-relevant criteria and human judgment rather than used as a single mechanical pass-fail threshold.
Norm-Referenced vs Ipsative Scoring
A common source of confusion in organizational assessment is the difference between norm-referenced scoring and ipsative scoring. Both produce numbers, but they are fundamentally different in what those numbers mean.
Norm-referenced scores place an individual's result on a scale that is calibrated against a population distribution. An individual's score on any one dimension is independent of their score on any other dimension, for example, high conscientiousness does not force low extraversion.
Ipsative scoring, used by instruments such as DISC, produces scores that are relative within the individual rather than relative to a population. Because ipsative formats ask respondents to distribute a fixed number of points across dimensions, a high score on one dimension necessarily produces lower scores on others. Ipsative scores cannot be used to benchmark an individual against a population norm, because the scores have no population-level reference point. This is a fundamental limitation for selection and development purposes. A high Dominance score in DISC means the individual chose Dominance responses over Influence, Steadiness, and Conscientiousness responses, not that they score higher on dominance-related traits than 70% of working adults.
How Deeper Signals Approaches Norming
Our assessments are normed against a global database of more than 150,000 working adults, with data collected across North America, Europe, South America, Africa, and Asia. This scale and breadth exceed what many legacy instruments rely on, where norms are often built from a few thousand respondents drawn from one or two countries and re-normed only rarely. Genuine cross-regional coverage at this scale means percentile estimates remain stable at the extremes of the distribution and within subgroups, which is precisely where smaller norms become unreliable.
The normative database is continuously updated as new data are collected, ensuring that percentile interpretations reflect the current distribution of scores in the relevant working-adult population rather than a snapshot from a single historical collection period. Currency is treated as an ongoing process, not a one-time event.
This approach also supports fair and defensible decisions. A large, current, and globally representative norm base reduces the risk that an ill-fitting reference population introduces spurious subgroup differences, which is part of what makes assessment-based decisions defensible under frameworks such as the Uniform Guidelines and the SIOP Principles. Norm-referenced scoring is also one of the safeguards that makes Deeper Signals assessments more resistant to response distortion: because faking tends to be relatively uniform across candidates, comparing each person against a relevant norm helps preserve the rank ordering that predictive validity depends on (see our companion article on faking).
Every score in a our assessments is expressed as a percentile relative to this global norm group, alongside an interpretive band that communicates whether the score represents a low, mid-range, or high position on the dimension. This is a deliberate design choice: pairing a precise percentile with a plain-language band makes the feedback meaningful to respondents without psychometric training while keeping interpretation anchored to a clearly defined reference population, which is exactly what Kline (1994) identifies as the purpose of norming.
Frequently Asked Questions
Does a normed assessment measure something more accurately than an unnormed one?
Norming does not improve the accuracy of measurement. Reliability and validity determine how accurately an instrument measures its intended construct. Norming determines whether scores can be interpreted meaningfully relative to a reference population. An unnormed assessment may measure its construct reliably and validly but cannot produce population-relative interpretations. Both measurement quality and norming are required for professional use.
What is a percentile score?
A percentile score indicates the percentage of the norm group that scored at or below a given raw score. A percentile of 70 means the individual scored at or above 70% of the norm group on that dimension. Percentile scores are the most intuitive norm-referenced format for non-specialist audiences.
What is a sten score?
A sten (standard ten) score is a standardized norm-referenced score expressed on a scale of 1 to 10, with a mean of 5.5, so the average of the norm group falls between stens 5 and 6. Each sten band is half a standard deviation wide. Sten scores are common in occupational personality assessments because they divide the distribution into ten bands that are easy to interpret and communicate. Scores of 1 to 3 represent the lower range, 4 to 7 the mid-range, and 8 to 10 the upper range relative to the norm group.
Can I use norms developed in one country for candidates in another?
Not without caution. Score distributions on personality assessments can vary across national populations, and applying norms developed in one country to candidates from another may produce systematically biased percentile interpretations. The Standards (AERA, APA, & NCME, 2014) require that cross-cultural use of norms be evaluated through measurement invariance studies that test whether the instrument functions equivalently across the populations in question.
How do I know if an assessment's norms are appropriate for my organization?
Request the technical manual and examine the norm group description: the size, demographic composition, occupational distribution, national origin, and date of collection. Then evaluate whether that population is sufficiently similar to your own candidate or employee population to produce meaningful percentile interpretations. If it is not, contact the provider and ask whether role-specific or industry-specific norms are available.
When should I use a global norm versus a role or region-specific norm?
A large, representative global norm is usually the right default for general selection and development, because it provides a broad and stable working-adult reference point. A more specific norm becomes worth requesting when you are assessing a population that differs systematically from the general working population in ways that matter for the decision, for example, senior executives, a single national population where cultural response differences are a concern, or a high-volume role where you want to benchmark against incumbents in that exact job. The practical test is whether a more specific norm would change interpretation enough to change decisions. If it would, ask the provider whether role, industry, or region-specific norms are available and how large and current they are.
Last reviewed by Dr. Reece Akhtar — June 2026
References
Kline, P. (1994). The Handbook of Psychological Testing. London: Routledge.
American Educational Research Association, American Psychological Association, & National Council on Measurement in Education. (2014). Standards for Educational and Psychological Testing. Washington, DC: American Educational Research Association.
Society for Industrial and Organizational Psychology. (2018). Principles for the Validation and Use of Personnel Selection Procedures (5th ed.). Bowling Green, OH: SIOP.
Equal Employment Opportunity Commission. (1978). Uniform Guidelines on Employee Selection Procedures. 29 C.F.R. Part 1607.


