Are Personality Tests Scientifically Valid?
Yes, with important qualifications. Personality assessments built on the Five Factor Model have meaningful, peer-reviewed validity evidence for predicting job performance, leadership potential, and team behavior. Understanding what the research actually shows, including where the evidence is stronger and where it is more limited, is what responsible use of personality data requires.
The Foundational Evidence
The modern case for personality testing in organizational settings rests primarily on meta-analytic research, i.e. studies that aggregate results across many individual studies to produce more stable estimates of validity.
Barrick and Mount (1991) conducted the most influential of these, reviewing 117 studies involving more than 23,000 participants across five occupational groups. They examined how each of the Big Five dimensions predicted job performance and training proficiency across managerial, professional, police, skilled, and semi-skilled roles. Their central finding was that Conscientiousness, the tendency to be organized, reliable, and goal-oriented, produced consistent positive validity across all occupational groups, with a corrected mean validity coefficient of .22 for job performance. No other Big Five dimension showed this consistency, though Openness to Experience predicted training proficiency across groups.
Salgado (1997) extended this evidence to European samples, conducting a meta-analysis of studies from across European Community member states. His findings confirmed that Conscientiousness (corrected r = .25) and Emotional Stability (corrected r = .19) predicted job performance across national contexts, providing cross-cultural evidence that the personality-performance relationship established in North American research is not culturally specific.
Tett, Jackson, and Rothstein (1991) reached higher overall estimates in a parallel meta-analysis, reporting a mean corrected validity of .24 across Big Five dimensions for job performance. The difference between Tett et al.'s estimates and those of Barrick and Mount reflects methodological differences rather than contradictory findings. Tett et al. used a confirmatory strategy that tested theoretically derived hypotheses about which traits should predict performance in which contexts, producing higher estimates when traits were matched to relevant criteria. This distinction matters because it points toward something the evidence consistently supports: personality validity is higher when the right trait is matched to the right criterion rather than when all traits are tested against all criteria indiscriminately.
What the Evidence Shows, and What It Does Not
Taken together, the meta-analytic evidence supports several specific conclusions.
Conscientiousness is the most consistently valid personality predictor of job performance across roles, industries, and cultures (Barrick & Mount, 1991; Salgado, 1997). Its validity is meaningful, though modest in absolute terms, and it does not approach the predictive validity of general cognitive ability (Schmidt & Hunter, 1998). Emotional Stability is a consistent predictor of job satisfaction and performs well under demanding conditions. Extraversion predicts performance in roles requiring social interaction, particularly sales and management. Agreeableness predicts performance in roles requiring teamwork and interpersonal cooperation. Openness to Experience predicts training success and performance in creative and innovative roles.
The personality-performance relationship is strongest when personality measures are used in combination with cognitive ability measures rather than as a standalone selection method. Sackett, Zhang, Berry, and Lievens (2022), the current reference for corrected validity estimates across selection methods, confirmed that predictive validity is substantially higher when personality and cognitive ability assessments are used together, because the two constructs capture independent contributions to performance. Neither alone produces as accurate a prediction as the combination.
The Scientific Debate
The validity evidence for personality testing has not gone unchallenged, and practitioners need to know that a genuine scientific debate exists.
Morgeson, Campion, Dipboye, Hollenbeck, Murphy, and Schmitt (2007) published an influential critique of personality testing in personnel selection, raising five substantive concerns: that validity coefficients are low in absolute terms; that faking and social desirability effects reduce the quality of self-report data; that personality constructs are too broadly defined to produce specific, job-relevant predictions; that insufficient job analysis is conducted before personality measures are deployed; and that criterion contamination inflates observed validity estimates.
These are serious criticisms and they reflect real limitations of personality testing as typically practiced. The critique does not argue that personality data has no validity. Morgeson et al. (2007) acknowledge the accumulated meta-analytic evidence in their work. The argument is that the validity is lower than proponents claim, that it is inconsistently obtained, and that the practical conditions required to produce it are often absent in real selection contexts.
Ones, Dilchert, Viswesvaran, and Judge (2007) published a direct rebuttal to Morgeson et al. in the same issue of Personnel Psychology. Their response addressed each point: corrected validity estimates are meaningful for selection decisions even when modest in absolute terms; faking effects are limited and detectable; broad personality traits predict broad criteria, which is the appropriate level of abstraction for most organizational outcomes; and the standard for how much job analysis is required before using personality measures is itself contested. The Morgeson-Ones exchange represents the field's most substantive recent debate about personality testing.
The most defensible position is that personality tests built on the Five Factor Model have genuine, replicable validity evidence that justifies their use in talent decisions, provided they are used as part of a multi-measure battery, matched to job-relevant criteria, selected for their documented psychometric properties rather than brand recognition, and interpreted with appropriate epistemic humility about the precision of prediction personality data can support.
What Makes a Personality Test Scientifically Valid?
Understanding whether a specific personality instrument is scientifically valid requires examining its technical documentation rather than assuming all personality tests are equivalent. A scientifically valid psychometric test meets several specific criteria.
It reports internal consistency (Cronbach's alpha) coefficients at or above 0.70 for each scale. It demonstrates test-retest reliability across meaningful time intervals. It reports criterion validity against job performance or other relevant outcome measures, not just convergent validity against other personality scales. It has been normed on a large, representative population so that scores can be interpreted relative to a meaningful reference group. And it has been tested for adverse impact across gender, age, and ethnic groups, with analyses available for review before deployment.
Instruments that lack any of these elements should not be described as scientifically validated personality assessments, regardless of how widely they are used or how intuitively appealing their output appears.
How Deeper Signals Approaches This
At Deeper Signals, the Core Drivers Diagnostic was developed with the specific intent of meeting the technical standards that distinguish validated from unvalidated personality instruments. Its development followed the psychometric evidence: built on the Big Five framework, developed using Genetic Algorithms across 50,000+ working adults to maximize convergent validity, validated against the NEO PI-R and Hogan Personality Inventory, and tested for adverse impact across gender, age, and ethnicity with all ratios at or above the professionally required 0.80 threshold.
The Core Drivers Diagnostic reports reliability coefficients between .69 and .82 across its six scales and demonstrates criterion validity against job performance, work engagement, and counterproductive work behaviors. It does not claim to predict performance with a precision the research cannot support. The research on personality testing supports its use as one validated instrument in a broader assessment battery and that is exactly how Deeper Signals designs its use.
Frequently Asked Questions
Do personality tests predict job performance?
Yes, with important qualifications. Big Five-based personality tests show consistent positive validity for predicting job performance, particularly for conscientiousness, which is valid across virtually all occupational groups (Barrick & Mount, 1991; Salgado, 1997). Validity is higher when the right trait is matched to the right criterion and when personality data is combined with cognitive ability measures.
How strong is the validity of personality tests?
Corrected validity coefficients for personality dimensions typically range from approximately .15 to .25 for job performance outcomes in meta-analytic research (Barrick & Mount, 1991; Tett et al., 1991). These are meaningful for selection purposes but modest in absolute terms. Personality validity is lower than cognitive ability validity and is best realized when multiple validated instruments are used in combination.
Do personality tests have a scientific debate behind them?
Yes. Morgeson et al. (2007) raised substantive criticisms of personality testing in selection, including concerns about low validity, faking, and insufficient job analysis. Ones et al. (2007) published a direct rebuttal. Both papers appeared in the same issue of Personnel Psychology and together represent the field's most substantive recent exchange on this question. The consensus position is that validated Big Five instruments have meaningful, replicable validity, but that the conditions for realizing that validity require careful instrument selection and appropriate use.
Can personality tests be faked?
Yes, applicants can inflate their responses on personality tests, particularly on Conscientiousness and Emotional Stability scales. However, meta-analytic research shows that faking effects are limited in magnitude and that validity holds even in applicant samples where some inflation is likely. Well-designed instruments include response consistency checks that can detect systematic distortion, leverage adaptive item banks for test security, and additional algorithms to detect irregular responses. Many of these techniques are employed by Deeper Signals to ensure fair and honest assessments.
Which personality dimension best predicts job performance?
Conscientiousness is the most consistently valid personality predictor of job performance across occupational groups, industries, and cultures (Barrick & Mount, 1991; Salgado, 1997). Emotional Stability is the strongest predictor of job satisfaction. Other dimensions show validity in specific contexts: Extraversion for sales and leadership roles, Agreeableness for teamwork-intensive roles, Openness for training and creative roles.
Last reviewed by Dr. Reece Akhtar — June 2026
References
Barrick, M. R., & Mount, M. K. (1991). The Big Five personality dimensions and job performance: A meta-analysis. Personnel Psychology, 44(1), 1–26.
Tett, R. P., Jackson, D. N., & Rothstein, M. (1991). Personality measures as predictors of job performance: A meta-analytic review. Personnel Psychology, 44(4), 703–742.
Salgado, J. F. (1997). The five-factor model of personality and job performance in the European Community. Journal of Applied Psychology, 82(1), 30–43.
Morgeson, F. P., Campion, M. A., Dipboye, R. L., Hollenbeck, J. R., Murphy, K., & Schmitt, N. (2007). Reconsidering the use of personality tests in personnel selection contexts. Personnel Psychology, 60(4), 683–729.
Ones, D. S., Dilchert, S., Viswesvaran, C., & Judge, T. A. (2007). In support of personality assessment in organizational settings. Personnel Psychology, 60(4), 995–1027.
Schmidt, F. L., & Hunter, J. E. (1998). The validity and utility of selection methods in personnel psychology: Practical and theoretical implications of 85 years of research findings. Psychological Bulletin, 124(2), 262–274. https://doi.org/10.1037/0033-2909.124.2.262
Sackett, P. R., Zhang, C., Berry, C. M., & Lievens, F. (2022). Revisiting meta-analytic estimates of validity in personnel selection: Addressing systematic overcorrection for restriction of range. Journal of Applied Psychology, 107(11), 2040–2068. https://doi.org/10.1037/apl0000537


