All guides

How to Choose a Personality Assessment Tool

Author
Dr. Reece Akhtar
CEO and Co-founder at Deeper Signals
Last reviewed
06/2026

Most organizations choose personality assessment tools based on brand familiarity, ease of use, or the recommendation of a consultant rather than on the evidence that actually determines whether a tool will improve talent decisions. The result is widespread use of instruments that have not been validated against job performance, lack published reliability data, or produce adverse impact that organizations have never measured. Choosing well requires understanding a small number of specific, verifiable criteria and applying them consistently to every tool under consideration.

The Five Criteria That Matter

1. Criterion validity: does it predict job performance?

A personality assessment used in talent decisions must demonstrate criterion validity. It’s a documented statistical relationship between its scores and subsequent job performance outcomes. This is the most important criterion and the one most frequently absent from vendor documentation.

Ask for the technical manual and look for studies in which assessment scores predicted manager-rated or objective performance criteria in independent samples. Barrick and Mount (1991), in the foundational meta-analysis of personality and job performance, established that Big Five-based instruments have genuine criterion validity for predicting performance across occupational groups. This is the evidence standard any professionally used personality tool should meet.

A second, distinct form of validity to verify is construct validity. It’s evidence that the assessment measures what it claims to measure. This is typically demonstrated through documented correlations with well-established instruments, such as the NEO PI-R, Hogan Personality Inventory, or Big Five measures, placing the new instrument within a recognizable nomological network. Construct validity is necessary but not sufficient. It confirms the assessment measures personality meaningfully. Criterion validity confirms those measurements predict performance outcomes. Both should be present in the technical manual.

2. Reliability: does it measure consistently?

A reliable instrument produces consistent scores from the same person under similar conditions. Three reliability statistics should be requested from any vendor.

Cortina (1993) established that internal consistency (Cronbach's alpha) values above .70 are generally adequate for professional use, with values above .80 considered strong. Test-retest reliability above .70 across four-week intervals is the standard for instruments claiming to measure stable personality traits. The standard error of measurement (SEM) per scale indicates how much a single score might vary due to measurement imprecision. Smaller SEMs indicate more precise measurement. All three statistics should be reported in the technical manual for every scale in the instrument, not just in aggregate.

3. Normative data: can you interpret scores meaningfully?

A personality score has no meaning in isolation. A normed instrument situates each score relative to a reference population, allowing practitioners to understand where a candidate sits compared to relevant comparison groups. Kline (1994) establishes that adequate norming is a prerequisite for professional use of any psychological instrument.

Evaluate the norm group for three properties: size (large enough to produce stable estimates across the full range), representativeness (does the population match your candidate pool?), and currency (when were the norms collected?). An instrument normed on university students in one country and deployed with senior executives in another produces systematically misleading percentile interpretations.

4. Adverse impact: does it treat candidate groups equitably?

Any personality assessment used in hiring must have documented adverse impact analyses across gender, age, and ethnicity. The Standards for Educational and Psychological Testing (AERA, APA, & NCME, 2014) require this documentation for instruments used in consequential decisions. Sackett, Zhang, Berry, and Lievens (2022) confirmed that adverse impact profiles differ across selection methods and must be evaluated before deployment.

Request pass rates and adverse impact ratios by demographic group. All ratios should be at or above 0.80. A vendor who cannot provide this documentation has not conducted the analysis, which is itself a significant red flag.

5. Scientific foundation: what framework does it use?

The framework underlying an instrument determines the quality and interpretability of the data it produces. Instruments built on the Five Factor Model have the strongest and most replicable evidence base for predicting occupational outcomes (Barrick & Mount, 1991; Sackett et al., 2022). Type-based instruments including MBTI and DISC do not have comparable criterion validity evidence and should not be used for consequential talent decisions. Pittenger (2005) documented the specific psychometric limitations of MBTI in professional contexts, and the same concerns apply to other type-based tools.

Red Flags to Watch For

No published technical manual. A professionally designed instrument has a technical manual documenting its reliability, validity, norming methodology, and adverse impact analyses. If a vendor cannot provide one, treat the tool as unvalidated regardless of how polished the user experience appears.

Validity claims based on face validity or user satisfaction. "Our clients love it" and "candidates find it intuitive" are not validity claims. Validity is a statistical property of the relationship between scores and outcomes, and it must be documented empirically.

Type-based scoring. Instruments that assign people to fixed categories, such as INFJ or D-type, have not demonstrated the criterion validity required for professional selection use. The conversion of continuous data into discrete types loses information and reduces predictive accuracy.

Undisclosed norm groups. If a vendor cannot describe the size, demographic composition, and collection date of their normative sample, the percentile interpretations the instrument produces are uninterpretable.

No adverse impact data. The absence of adverse impact documentation is not evidence that the tool is unbiased. It is evidence that the vendor has not tested for bias.

Questions to Ask Every Vendor

Before signing a contract, ask for written answers to the following:

  • What is the criterion validity evidence for this instrument against job performance outcomes?
  • What are the internal consistency and test-retest reliability coefficients for each scale?
  • Describe the normative sample: size, demographic composition, occupational distribution, and date of collection.
  • Provide adverse impact analyses across gender, age, and ethnicity, including pass rates and Four-Fifths Rule calculations.
  • Which occupational groups and seniority levels has the instrument been validated for?
  • Is criterion validity evidence available in independent samples or only in the development sample?
  • If the assessment will be used across different countries or languages, can you provide completed measurement invariance results confirming the instrument functions equivalently across those populations?

A vendor who cannot answer all of these questions in writing is a vendor whose tool you cannot fully evaluate before deployment.

How Deeper Signals Approaches This

The Deeper Signals Suite of assessments meets each of the five criteria above. The Core Drivers Diagnostic demonstrates criterion validity against manager-rated task performance, contextual performance, work engagement, and counterproductive work behaviors. Internal consistency ranges from .69 to .82 across its six scales. Test-retest reliability is in the range of .68 to .77 across four-week intervals. The normative database covers 50,000+ working adults globally, with norms updated regularly. All adverse impact ratios are at or above the required 0.80 threshold across gender, age, and ethnicity.  The Core Drivers Diagnostic is built on the Five Factor Model, the framework with the strongest and most replicated evidence base for predicting work outcomes across cultures and occupational groups.

The technical manuals documenting all of this for the Core Drivers Diagnostic and other assessments are available on request before any purchasing decision. At Deeper Signals, providing complete documentation before contracting is a standard expectation, not an exceptional concession.

Frequently Asked Questions

How long should a technically adequate personality assessment take to complete?

Assessment length should be determined by what is needed to achieve adequate reliability, not by what feels comfortable. A well-designed short-form instrument can achieve reliability above .70 in 10 to 20 minutes. Instruments requiring 45 minutes or more are not automatically more valid and may create unnecessary candidate burden.

Is a free personality assessment ever appropriate for hiring?

Free instruments are rarely appropriate for consequential hiring decisions because they almost never have published criterion validity evidence, documented adverse impact analyses, or peer-reviewed normative data. They may be useful for low-stakes self-awareness purposes where predictive validity is not the goal.

Can I use different personality assessments for different roles?

Yes, and in some cases this is appropriate when role-specific validation data supports different instruments for different contexts. However, using different instruments across candidate groups for the same role creates non-comparable data and legal risk. Consistency within a selection process is a prerequisite for defensibility.

What is the difference between a personality assessment and a personality test?

In professional practice, the terms are often used interchangeably. "Psychometric test" is the broader professional term that implies adherence to technical standards including reliability, validity, and norming. Not every instrument labelled a "personality assessment" meets these standards, which is precisely why evaluating the underlying evidence is essential.

How do I know if a vendor's validity claims are credible?

Request the original validity studies, not a summary slide. Credible claims include sample sizes, independent validation samples, criterion measures used, and correlation coefficients. Be skeptical of validity claims without methodology, without independent replication, or based on user satisfaction rather than statistical prediction.

Last reviewed by Dr. Reece Akhtar — June 2026

References

Barrick, M. R., & Mount, M. K. (1991). The Big Five personality dimensions and job performance: A meta-analysis. Personnel Psychology, 44(1), 1–26.

Cortina, J. M. (1993). What is coefficient alpha? An examination of theory and applications. Journal of Applied Psychology, 78(1), 98–104. https://doi.org/10.1037/0021-9010.78.1.98

Kline, P. (1994). The Handbook of Psychological Testing. London: Routledge.

Sackett, P. R., Zhang, C., Berry, C. M., & Lievens, F. (2022). Revisiting meta-analytic estimates of validity in personnel selection: Addressing systematic overcorrection for restriction of range. Journal of Applied Psychology, 107(11), 2040–2068. https://doi.org/10.1037/apl0000537

American Educational Research Association, American Psychological Association, & National Council on Measurement in Education. (2014). Standards for Educational and Psychological Testing. Washington, DC: American Educational Research Association.

Pittenger, D. J. (2005). Cautionary comments regarding the Myers-Briggs Type Indicator. Consulting Psychology Journal: Practice and Research, 57(3), 210–221. https://doi.org/10.1037/1065-9293.57.3.210

Subscribe
Subscribe to the Deeper Signals newsletter
Thank you! Your submission has been received!
Please fill all fields before submiting the form.
Curious to learn more?

Schedule a call with Deeper Signals to understand how our assessments and feedback tools help people gain a deep awareness of their talents and reach their full potential. Underpinned by science and technology, we build talented people, leaders and companies.

  • Scalable and engaging assessment solutions
  • Measurable and predictive talent insights
  • Powered by technology and science that drives results
Let's talk!
  • Scalable interventions for growth
  • Measurable data, insights and outcomes for high performance
  • Proven scientific expertise that links results to outcomes
Thank you!
Would you like to schedule a meeting?
Please fill all fields before submitting the form.