How to Reduce Bias in Hiring with Assessments
Unstructured hiring is one of the most biased processes in organizational life. When interviewers evaluate candidates without standardized questions, consistent criteria, or documented scoring, their judgments are shaped by affinity bias, halo effects, demographic stereotypes, and first-impression anchoring in ways that are difficult to detect and impossible to audit. Structured, validated assessments reduce these sources of error, but they do not eliminate bias automatically. Using assessments to genuinely reduce bias in hiring requires choosing the right instruments, deploying them correctly, and understanding the evidence honestly.
Why Informal Hiring Produces Biased Outcomes
The evidence that unstructured hiring introduces bias is extensive and consistent. Sackett and Lievens (2008), in their comprehensive review of personnel selection research in the Annual Review of Psychology, documented that unstructured interviews, which are the default selection method in most organizations, are among the most susceptible to interviewer bias of any commonly used selection procedure. Interviewers evaluate candidates differently depending on the interviewer's own characteristics, the order in which candidates are seen, and demographic factors that are legally irrelevant to job performance. The subjectivity that gives unstructured interviews their perceived flexibility is the same property that makes them unreliable and biased.
Structured assessments replace subjective impression with standardized, comparable data. Every candidate completes the same instrument under the same conditions, scored against the same normative database. This consistency is the foundation of bias reduction. It ensures that irrelevant variation between candidates (the interviewer's mood, the time of day, the candidate's social similarity to the evaluator) does not influence the evaluation in the way it does in unstructured processes.
What Assessment Can and Cannot Do About Bias
The relationship between validated assessment and bias is more nuanced than either "assessments are objective so they are bias-free" or "assessments replicate existing inequalities." Both of these positions misrepresent the evidence.
Schmidt and Hunter (1998), whose synthesis of 85 years of selection research remains the foundational reference in this field, documented that selection methods differ substantially in both their predictive validity and their adverse impact profiles. Cognitive ability tests produce the highest validity for predicting job performance but also show the largest mean score differences between demographic groups. Personality assessments show meaningful validity with smaller group differences. Structured interviews show strong validity with relatively small group differences.
Gottfredson (1986) documented the societal and organizational consequences of the g factor, which stands for general cognitive ability, in employment, establishing that the validity-diversity trade-off is a real empirical phenomenon rather than a theoretical concern. High-validity selection methods do not automatically minimize group differences, and organizations face a genuine trade-off between maximizing predictive accuracy and minimizing differential impact across demographic groups. Acknowledging this trade-off honestly is the starting point for responsible assessment design.
Berry, Lievens, Zhang, and Sackett (2024), in the most current meta-analytic update of the personnel selection validity matrix, confirmed that these patterns hold in contemporary data. Cognitive ability remains high-validity with larger subgroup differences, while combinations of methods can reduce adverse impact while maintaining meaningful validity.
How to Use Assessments to Reduce Bias
Step 1: Choose instruments tested for adverse impact. Every assessment used in hiring should have published adverse impact analyses across gender, age, and ethnicity before deployment. Requesting this documentation from vendors is the minimum standard for responsible procurement. The Uniform Guidelines on Employee Selection Procedures (1978) require that any selection procedure with adverse impact be validated for the job in question.
Step 2: Use a multi-method battery. A pre-employment assessment battery combining cognitive ability, personality, and structured competency-based interviews produces higher validity than any single method while distributing adverse impact across instruments with different subgroup difference profiles. Sackett and Lievens (2008) established that combinations of selection methods can reduce adverse impact relative to using a high-validity cognitive ability test as the sole criterion, without sacrificing predictive accuracy.
Step 3: Set evidence-based cutoff scores. Applying a fixed minimum score on a cognitive ability test as the sole screen is the highest-risk approach for adverse impact. Using scores as one weighted input in a composite scoring model, combined with personality, values, or structured interview data, preserves predictive validity while reducing the probability that any single instrument's subgroup differences drive the entire selection outcome. Validate cutoff scores against local performance data before deploying them at scale.
Step 4: Document everything. Bias reduction is a recurring audit process. Document which instruments are used at each selection stage, what the adverse impact ratios are for your specific applicant population, and what actions are taken when ratios fall below the 0.80 threshold required by the Uniform Guidelines. This documentation is both a legal requirement and a genuine diagnostic tool for identifying where the process is producing inequitable outcomes.
Common Mistakes
Assuming validation means bias-free. A validated assessment has been tested for predictive validity. It predicts job performance. That is a different property from adverse impact, which measures whether the assessment produces different outcomes across demographic groups. An assessment can be both valid and produce adverse impact. Both properties must be documented and reviewed independently.
Using cognitive ability as the sole screen. Cognitive ability tests have the strongest validity evidence of any single selection method, but they also show the largest group mean score differences. Using cognitive ability as a sole cutoff maximizes predictive accuracy and maximizes adverse impact simultaneously. The appropriate use is as one component of a structured multi-method battery.
Treating structured assessment as a substitute for bias training. Assessments standardize one part of the selection process, which is the measurement stage. Bias can still enter at the job description stage, at the sourcing stage, during offer negotiation, and in promotion decisions. Assessment reduces bias where it is deployed. It does not address bias in the rest of the talent lifecycle.
How Deeper Signals Approaches This
Bias minimization is a design requirement for every assessment on the Deeper Signals platform, not a property of any single instrument. Every assessment, including the Core Drivers Diagnostic (personality), the Core Values Diagnostic (values and motivation), the Core Reasoning assessment (cognitive ability), and the competencies of the Deeper Signals Capability Model, is tested for adverse impact across gender, age, and ethnicity. In the analyses available, adverse impact ratios sit at or above the four-fifths (0.80) screening threshold using recommended cutoff scores, and Deeper Signals supports clients in monitoring impact within their own deployments.
This is achieved through deliberate design. The techniques used across our assessments include: differential item functioning (DIF) analysis to detect and remove items that perform differently across demographic groups; forced-choice, desirability-matched item formats in the personality assessment that reduce response distortion and the cultural legibility of "better" answers; behaviorally framed item content tied to observable actions rather than evaluative labels; norm-referenced scoring against a large, globally representative norm base so that interpretation is anchored to a relevant population; and multi-trait composite scoring, with all custom scoring algorithms tested for adverse impact during development.
Frequently Asked Questions
Do personality assessments produce less adverse impact than cognitive ability tests?
Generally, yes. Personality assessments built on the Big Five show smaller mean score differences between demographic groups than cognitive ability tests across most demographic comparisons. This makes personality assessment a valuable complement to cognitive ability in a multi-method battery, adding validity while moderating the overall battery's adverse impact profile (Sackett & Lievens, 2008).
What is the 4/5ths rule in adverse impact analysis?
The 4/5ths rule, from the Uniform Guidelines on Employee Selection Procedures, states that if the selection rate for any protected group is less than 80% of the rate for the highest-scoring group, adverse impact is indicated. An adverse impact ratio below 0.80 does not automatically mean a tool is biased. It triggers a requirement to investigate and document.
Can structured interviews reduce bias compared to unstructured ones?
Yes. Structured interviews standardize questions, scoring criteria, and evaluation procedures across all candidates, significantly reducing the influence of interviewer-specific biases and demographic stereotypes. The validity advantage of structured over unstructured interviews is well documented (McDaniel et al., 1994), and their more favorable adverse impact profile makes them a strong component of a bias-aware selection process.
What should I do if my assessment produces adverse impact in my organization?
First, verify the calculation using your own applicant data. Adverse impact ratios can vary by population and context. Second, review whether cutoff scores can be adjusted to reduce impact while maintaining adequate validity. Third, consider adding a complementary instrument with a different adverse impact profile to the battery. Fourth, document all findings and actions taken. And fifth, consult with the assessment provider and legal counsel before making changes to a high-stakes selection process.
Last reviewed by Dr. Reece Akhtar — June 2026
References
Gottfredson, L. S. (1986). Societal consequences of the g factor in employment. Journal of Vocational Behavior, 29, 379–410.
Schmidt, F. L., & Hunter, J. E. (1998). The validity and utility of selection methods in personnel psychology: Practical and theoretical implications of 85 years of research findings. Psychological Bulletin, 124(2), 262–274. https://doi.org/10.1037/0033-2909.124.2.262
Sackett, P. R., & Lievens, F. (2008). Personnel selection. Annual Review of Psychology, 59, 419–450. https://doi.org/10.1146/annurev.psych.57.102904.190200
Berry, C. M., Lievens, F., Zhang, C., & Sackett, P. R. (2024). Insights from an updated personnel selection meta-analytic matrix. Journal of Applied Psychology, 109(10), 1611–1634. https://doi.org/10.1037/apl0001203
McDaniel, M. A., Whetzel, D. L., Schmidt, F. L., & Maurer, S. D. (1994). The validity of employment interviews: A comprehensive review and meta-analysis. Journal of Applied Psychology, 79(4), 599–616.


