Is AI CV Screening Actually Fair?
Not automatically, but it can be, and the honest answer is more useful than either extreme in the current debate. AI-driven CV screening does not become fair simply because it is more consistent than a human reviewer, and it does not become inherently biased simply because it is powered by machine learning.
What AI CV Screening Actually Does
AI-driven CV screening tools typically use natural language processing to extract information from resumes and applications, such as skills, experience, and qualifications, and then score or rank candidates against criteria the model has learned to associate with a good fit. These tools are attractive because they can process far more applications than a human recruiter could review manually, and they can apply the same criteria consistently across every application, at least in theory.
The fairness question is not whether these tools are consistent. Consistency alone does not guarantee fairness. A model that applies a biased criterion consistently across every candidate is not fairer than a human evaluator who applies that same bias inconsistently. It may actually be worse, since the bias now affects every candidate the same way, at much greater scale.
There is also a second layer to consider. Boyce, Hickman, and Boyce (2026) report that an estimated 40 percent of applicants already use AI to customize their resumes. This means AI increasingly sits on both sides of the process, shaping the documents that AI screeners then read.
What the Research Actually Shows
Zhang et al. (2023) demonstrated something genuinely important. Subgroup score differences in machine learning-based selection are not fixed or inevitable. Their research showed that models explicitly designed with a dual objective, maximizing predictive accuracy while simultaneously minimizing group differences, can reduce subgroup gaps relative to models optimized for prediction alone. Related work from the same research effort found that deliberately oversampling high-scoring candidates from underrepresented groups in the training data can also reduce adverse impact in the resulting model.
Both of these are active, deliberate design choices, not something that happens automatically when an organization adopts machine learning. Left to optimize purely for predictive accuracy on historical data, a model has no inherent reason to reduce subgroup differences, and depending on what patterns exist in that historical data, it may just as easily reproduce or amplify them.
That risk is not hypothetical when AI is used off the shelf. In a review of AI in selection, Boyce, Hickman, and Boyce (2026) report that LLMs tend to score resumes carrying a minority-group indicator lower than otherwise-equivalent resumes. Examples include a common minority name or a disability-related award, and the simple prompts most accessible to everyday users are especially likely to produce biased results.
This connects directly to the broader picture from Hunkenschroer and Luetge's (2022) review of AI recruiting ethics, covered in more depth in our post on algorithmic bias. The actual effect of AI on hiring fairness is genuinely unresolved in the research and depends heavily on implementation choices, not on the mere presence of AI in the process.
The Trade-Off Worth Understanding
Building fairness into a model as an explicit objective is not free. When a model is optimized to reduce subgroup differences alongside maximizing prediction, it can produce what researchers call differential prediction, meaning the model's relationship to the outcome it is trying to predict is not identical across every group.
This is not necessarily a reason to avoid fairness-aware design. It is a reason to be transparent about the trade-off rather than pretending it does not exist. Organizations deploying AI CV screening should understand that fairness and pure predictive optimization are not always perfectly aligned, and that a responsible vendor should be able to explain exactly how their model balances the two, rather than claiming to have maximized both without any trade-off at all.
What Actually Determines Whether AI CV Screening Is Fair
Whether fairness was an explicit design objective. Ask any vendor directly whether subgroup differences were measured and addressed during model development, and what specific techniques, such as those Zhang et al. (2023) studied, were used to do so.
What the model was trained to predict. A model trained on which candidates were previously interviewed or hired will tend to reproduce the patterns in that historical data. A model trained on a more carefully defined, job-relevant outcome has a better chance of avoiding this.
Whether subgroup differences are tested on an ongoing basis, not just at launch. Adverse impact can emerge or shift as a model is used on new populations over time, even if it appeared fair in initial testing.
Whether the underlying features driving the model's decisions are actually job-relevant. A model that scores resumes well partly based on features correlated with protected characteristics, such as certain schools, neighborhoods, or gaps in employment, is not made fair simply by having a good overall predictive accuracy score.
How Deeper Signals Approaches This
At Deeper Signals, candidate evaluation is grounded in validated, transparent psychometric assessment rather than machine learning models trained to score unstructured resume text.
The Core Drivers Diagnostic and Core Values Diagnostic are tested for adverse impact across gender, age, and ethnicity as a standard part of their validation, with all ratios documented and available for review. Where the Deeper Signals platform does use job-matching algorithms, such as mapping candidates to role profiles, the principle Zhang et al. (2023) demonstrated applies directly. Fairness has to be tested and designed for, not assumed. This is why Deeper Signals' own research on AI-assisted job profiling included a formal adverse impact analysis across demographic groups before deployment.
Frequently Asked Questions
Is AI CV screening always more biased than human screening?
No, and it is not automatically fairer either. Zhang et al. (2023) found that machine learning models can reduce subgroup differences relative to models optimized purely for prediction, but this requires fairness to be an explicit design objective, not an assumption.
What is differential prediction, and why does it matter?
Differential prediction means a model's relationship to the outcome it predicts is not identical across demographic groups. It can arise when fairness-aware design techniques are used, and organizations should understand this trade-off rather than assuming a model has perfectly optimized both fairness and accuracy simultaneously.
Can oversampling underrepresented candidates in training data actually reduce bias?
Research associated with Zhang et al.'s (2023) special issue found that oversampling high-scoring candidates from underrepresented groups in training data can reduce adverse impact in the resulting model, though this is one specific technique among several and is not a universal fix for every context.
What should I ask an AI CV screening vendor about fairness?
Ask whether subgroup differences were measured and explicitly addressed during model development, what specific techniques were used, how the model performs on your organization's own applicant data, and how often fairness is re-tested after deployment.
Does removing demographic information from resumes before AI screening solve the fairness problem?
Not entirely. Machine learning models can learn to infer demographic information indirectly through proxy variables, such as school names, hobbies, or employment gaps, even when explicit demographic fields are removed. Testing the model's actual outputs for subgroup differences is more reliable than assuming removed fields eliminate the risk.
Last reviewed by Dr. Reece Akhtar — June 2026
References
Boyce, A. S., Hickman, L., & Boyce, C. E. (2026). The future of selection enabled by artificial intelligence. In N. Schmitt & A. M. Ryan (Eds.), The Oxford handbook of personnel assessment and selection (2nd ed.). Oxford University Press. https://doi.org/10.1093/9780197809013.003.0018
Zhang, N., Wang, M., Xu, H., Koenig, N., Hickman, L., Kuruzovich, J., Ng, V., Arhin, K., Wilson, D., Song, Q. C., Tang, C., Alexander, L., & Kim, Y. (2023). Reducing subgroup differences in personnel selection through the application of machine learning. Personnel Psychology, 76(4), 1125–1159. https://doi.org/10.1111/peps.12593
Hunkenschroer, A. L., & Luetge, C. (2022). Ethics of AI-enabled recruiting and selection: A review and research agenda. Journal of Business Ethics, 178(4), 977–1007. https://doi.org/10.1007/s10551-022-05049-6


