What Is Predictive Validity in Hiring?
Predictive validity is the statistical relationship between a score on a selection assessment and a subsequent measure of job performance. In hiring, it answers a specific and consequential question: does the information collected about a candidate before hire actually predict how well they will perform once in the role? A selection method with high predictive validity produces scores that correlate strongly with later job performance. A method with low predictive validity produces scores that are largely unrelated to how well the person goes on to perform. Predictive validity is the single most important psychometric property a pre-employment assessment can have, and it is the primary basis on which hiring methods should be evaluated.
What Predictive Validity Means, and How It Is Measured
Predictive validity is commonly expressed as a correlation coefficient, typically denoted r, ranging from 0 to 1. A coefficient of 0 indicates that scores on the assessment are entirely unrelated to subsequent job performance, meaning knowing a candidate's score tells you nothing about how they will perform. A coefficient of 1 indicates perfect prediction, meaning knowing the score would allow perfect forecasting of performance. In practice, no selection method achieves perfect prediction. Real-world validity coefficients for established selection methods range from approximately .20 to .55 after statistical corrections (Schmidt & Hunter, 1998; Sackett et al., 2022).
Two statistical corrections are routinely applied to raw validity estimates to produce more accurate population estimates. Range restriction correction accounts for the fact that organizations only see the performance of people they hired, not the full range of applicants, which compresses the observed correlation. Criterion reliability correction accounts for the fact that job performance measures are themselves imperfect and contain measurement error. Both corrections increase the observed coefficient, producing estimates that better reflect the true relationship between assessment scores and job performance in the full applicant population.
The debate about how much to correct, and which correction procedures to apply, is an active one in the scientific literature. Sackett, Zhang, Berry, and Lievens (2022) revisited the meta-analytic validity estimates across selection methods and identified that some historical estimates were overcorrected using procedures that inflated validity coefficients beyond what the data supported. Their corrected estimates are the current reference point for practitioners evaluating the evidence base for different selection methods.
Which Methods Predict Best According to Evidence?
Hunter and Hunter (1984) conducted the foundational meta-analytic comparison of selection method validity, examining how well different predictors, such as cognitive ability tests, personality measures, interviews, reference checks, biographical data, and others, predicted job performance and training success. Their central finding established cognitive ability as the strongest single predictor of job performance, outperforming all alternative methods examined.
Schmidt and Hunter (1998) extended this synthesis across 85 years of accumulated personnel selection research, providing the most comprehensive ranking of selection methods by predictive validity available at the time. Their analysis produced a hierarchy of methods that remains influential in the field. General mental ability (GMA) tests showed a corrected validity of approximately .51 for job performance when used alone. Structured interviews produced approximately .51 when used alone, and the combination of GMA with a structured interview reached approximately .63. The combination of GMA with an integrity test reached approximately .65, the highest two-predictor combination Schmidt and Hunter examined. Unstructured interviews produced approximately .38. Personality assessments based on the Big Five produced approximately .31 for Conscientiousness dimension used alone, with the personality-GMA combination producing substantially higher estimates.
Salgado, Anderson, Moscoso, Bertua, and De Fruyt (2003) confirmed that these validity findings generalize beyond the North American research base. Their meta-analysis of GMA validity across European Community member states found that cognitive ability predicted both job performance and training success across national contexts with validity estimates closely comparable to North American figures, establishing that predictive validity is not a culturally specific phenomenon.
What Recent Evidence Tells Us
Sackett, Zhang, Berry, and Lievens (2022) revisited the validity estimates that have guided practitioners since Schmidt and Hunter (1998) and identified a systematic problem: many historical meta-analytic estimates applied range restriction corrections that assumed a level of restriction more severe than the data actually showed. The result was that several validity estimates were inflated, some meaningfully so.
After applying more appropriate corrections, Sackett et al. (2022) found that GMA validity for job performance was lower than Schmidt and Hunter's (1998) estimates, though cognitive ability remained among the highest-validity methods. Structured interviews showed higher validity than previously estimated in some analyses. Personality assessments showed validity that was broadly consistent with earlier estimates, maintaining their position as meaningful predictors, particularly when combined with cognitive ability.
The practical implication of Sackett et al. (2022) is not that selection science was wrong. It is that the field updated its estimates when better methods became available, which is how science is supposed to work. The hierarchy of methods remains largely consistent: cognitive ability and structured psychometric testing are among the most valid approaches; unstructured interviews and informal reference checks are among the weakest; and the strongest prediction is achieved by combining multiple validated instruments that capture independent contributions to performance.
Why Predictive Validity Matters for Hiring Practice
Predictive validity is not an abstract psychometric concept. It has direct organizational consequences. A selection method with higher predictive validity produces a better match between the candidate selected and the demands of the role, reducing the probability of early attrition, performance problems, and the cost of rehiring.
Schmidt and Hunter (1998) quantified this through utility analysis: the dollar value of selecting a higher-performing employee rather than an average one multiplies across the number of hires made and the tenure of those hires. Organizations that use high-validity selection methods systematically select better-performing employees than organizations relying on low-validity methods, and that difference compounds at scale.
The practical guidance from the evidence is consistent: build selection processes around the highest-validity methods available, such as cognitive ability assessments, structured interviews, and validated personality assessments, and use them in combination rather than relying on any single method alone (Hunter & Hunter, 1984; Schmidt & Hunter, 1998; Sackett et al., 2022).
How Deeper Signals Approaches Predictive Validity
At Deeper Signals, predictive validity is a design requirement for every assessment in the platform. The Core Drivers, Core Values and Core Reasoning assessments demonstrates criterion validity against job performance, work engagement, and counterproductive work behaviors.
Every Deeper Signals’ assessment is designed to be used in combination, reflecting the research consensus that predictive validity is maximized when multiple validated instruments capturing independent constructs are combined into a structured battery rather than used in isolation.
Frequently Asked Questions
What is the difference between predictive validity and construct validity?
Predictive validity measures whether assessment scores predict a future criterion, typically job performance. Construct validity measures whether an assessment actually measures the psychological construct it claims to measure, such as conscientiousness or general reasoning ability. Both are required for a professionally defensible assessment: an instrument needs to measure what it claims to measure and that measurement needs to predict the outcomes that matter.
What is a good predictive validity coefficient for a hiring assessment?
In selection research, corrected validity coefficients above .30 are generally considered practically meaningful for predicting job performance. Coefficients above .50 are strong. Most selection methods used in isolation fall between .20 and .55 after appropriate corrections. The highest predictive validity is achieved by combining multiple validated instruments (Schmidt & Hunter, 1998; Sackett et al., 2022).
Why do different studies report different validity figures for the same assessment?
Validity estimates vary because of differences in the job studied, the performance criteria used, the sample size, and the statistical corrections applied. Meta-analytic research aggregates estimates across studies to produce more stable figures. Sackett et al. (2022) identified that some historical estimates were inflated due to overcorrection for range restriction, which is why current reference figures differ from those published in earlier syntheses.
Does predictive validity change across job types?
Yes. Predictive validity is typically higher for complex jobs with significant learning demands, where general reasoning ability and dispositional factors play a larger role in performance outcomes. It is lower for highly routinized roles where performance is largely determined by procedural compliance rather than individual capability differences.
How do I evaluate whether an assessment provider's validity claims are credible?
Request the technical manual and examine whether it reports criterion validity, such as correlation between assessment scores and measured job performance, rather than only convergent or construct validity. Check the sample sizes and whether the validity studies used independent samples. Verify that the performance criterion was measured independently of the assessment scores.
Last reviewed by Dr. Reece Akhtar — June 2026
References
Hunter, J. E., & Hunter, R. F. (1984). Validity and utility of alternative predictors of job performance. Psychological Bulletin, 96(1), 72–98. https://doi.org/10.1037/0033-2909.96.1.72
Schmidt, F. L., & Hunter, J. E. (1998). The validity and utility of selection methods in personnel psychology: Practical and theoretical implications of 85 years of research findings. Psychological Bulletin, 124(2), 262–274. https://doi.org/10.1037/0033-2909.124.2.262
Salgado, J. F., Anderson, N., Moscoso, S., Bertua, C., & De Fruyt, F. (2003). International validity generalization of GMA and cognitive abilities: A European Community meta-analysis. Personnel Psychology, 56(3), 573–605. https://doi.org/10.1111/j.1744-6570.2003.tb00751.x
Sackett, P. R., Zhang, C., Berry, C. M., & Lievens, F. (2022). Revisiting meta-analytic estimates of validity in personnel selection: Addressing systematic overcorrection for restriction of range. Journal of Applied Psychology, 107(11), 2040–2068. https://doi.org/10.1037/apl0000537


