All guides

Can AI Replace Human Judgment in Hiring?

Author
Dr. Reece Akhtar
CEO and Co-founder at Deeper Signals
Last reviewed
06/2026

The evidence gives a specific and somewhat uncomfortable answer. Algorithmic combination of hiring data consistently outperforms human experts making the same decision through unstructured, holistic judgment. This is not a new finding produced by modern AI. It is a research question that personnel psychologists have studied since the 1950s, and the evidence has pointed in the same direction for decades. What AI has changed is the scale and sophistication of the algorithms available, not the underlying finding about which approach produces more accurate decisions.

The Question AI Inherited

Long before machine learning existed, personnel psychologists asked a specific question: when combining multiple pieces of information about a candidate, such as test scores, interview ratings, and reference checks, into a single hiring decision, is it better to have an expert weigh the evidence holistically, or to combine the same information using a fixed statistical formula?

This is known as the mechanical versus clinical prediction debate. Clinical prediction refers to an expert reviewing the available information and forming a holistic judgment. Mechanical prediction refers to combining the same information using a predetermined statistical rule, removing the expert's subjective weighting of the evidence at the final combination step. Modern AI-driven hiring tools are, in this framework, a sophisticated form of mechanical prediction.

What the Evidence Shows

Kuncel, Klieger, Connelly, and Ones (2013) conducted a comprehensive meta-analysis specifically comparing mechanical and clinical data combination across selection and admissions decisions. Their findings were striking in their consistency. Mechanical combination of data outperformed expert holistic judgment by more than 50% in predicting relevant outcomes, and this advantage held across the different types of decisions and data examined in their analysis.

The explanation for this finding is not that algorithms are smarter than experts. It is that mechanical combination is more consistent. Human experts, even highly trained ones, apply inconsistent weight to the same piece of evidence across different candidates, influenced by fatigue, mood, the order in which information is reviewed, and countless small situational factors that have nothing to do with the candidate's actual qualifications. A fixed formula applies the same weighting every time, eliminating this specific source of error.

Why This Finding Is Often Misinterpreted

The mechanical-versus-clinical finding is frequently misread as evidence that algorithms should replace human judgment entirely in hiring. This is a significant overreach of what the research actually supports.

Kuncel et al.'s (2013) finding is specifically about the combination step, the point at which multiple pieces of already-collected information are weighed together into a final decision. It says nothing about whether the underlying data being combined is itself valid, fair, or complete. An algorithm that mechanically combines biased or invalid inputs will reliably produce biased or invalid outputs. Mechanical combination improves consistency in weighing evidence. It does not manufacture good evidence out of poor inputs.

The finding also does not address who should decide what data is collected in the first place, how that data should be interpreted in context, or how to handle genuinely novel situations that a fixed formula was not designed to anticipate. These remain human responsibilities regardless of how the final combination step is performed.

AI Does Not Fix a Flawed Process, It Accelerates It

Dr. Luke Treglown, Deeper Signals' Director of AI and Assessment R&D, makes a related point. Using algorithms to identify and rank candidates is not a new idea, he notes, particularly in the context of AI-driven candidate recommendation. What AI has changed is the scale and speed at which this already-existing practice operates.

This creates a dynamic worth naming directly. Candidates increasingly use AI to write and optimize their resumes, while employers increasingly use AI to screen and filter those same resumes. Dr. Treglown describes this as an AI-versus-AI loop, in which candidates optimize for algorithms, algorithms evaluate optimized content, and it becomes genuinely unclear what is actually being assessed in the process.

The deeper implication is that AI does not resolve the core challenges that have always determined hiring quality: whether the data being used is accurate and job-relevant, whether it is fair and unbiased, and whether it actually predicts meaningful outcomes. Mechanical combination, whether performed by a simple formula or a sophisticated model, can scale a well-designed process just as easily as it can scale a flawed one. If the underlying data or design is poor, AI will reliably make the resulting decisions worse at greater speed and volume, not better.

What AI Cannot Yet Replace

Several specific functions in hiring remain, based on current evidence, better suited to human judgment than to algorithmic replacement.

Determining what to measure. Deciding which competencies, traits, or criteria actually matter for a specific role requires job analysis and organizational context that an algorithm does not independently generate. This is a human and organizational decision that precedes any algorithmic combination.

Interpreting genuinely novel or ambiguous situations. A candidate with an unusual career path, a nontraditional background, or a situation the underlying data was not designed to anticipate requires human judgment that a fixed formula, by definition, cannot flexibly provide.

Evaluating the fairness of the inputs themselves. An algorithm can consistently combine the data it is given, but determining whether that data reflects legitimate, job-relevant signal or encodes historical bias requires ongoing human oversight, testing, and accountability. When Stockdale, Hickman, and Liu (2026) had LLMs score employment interviews, the scores showed consistent subgroup differences, often larger than those in human ratings. Adverse impact must therefore be tested directly, not assumed away because a tool performs well overall.

Making the final consequential decision. Even where mechanical combination outperforms holistic judgment at weighing evidence, responsible organizations retain human accountability for the ultimate hiring decision, particularly given the legal and ethical stakes involved.

How to Combine Mechanical and Human Judgment Responsibly

The strongest approach, supported by the broader selection literature, uses mechanical combination for what it does well and preserves human judgment for what it does well, rather than treating the choice as all-or-nothing.

Collect data through validated, structured methods: validated psychometric assessments, structured interviews with predetermined scoring criteria, and job-relevant work samples. Combine that data using a validated, transparent formula rather than an opaque black-box algorithm, so the combination step benefits from mechanical consistency without sacrificing auditability. Reserve human judgment for interpreting context, evaluating genuinely novel situations, and taking final accountability for the decision.

Recent evidence on modern AI points in the same direction. Stockdale, Hickman, and Liu (2026) tested large language models scoring employment interviews and compared them to human raters. Overall, LLM and human interview scores were closely aligned, and a single human rating combined with one LLM rating reached criterion validity comparable to a panel of three human raters.

The practical pattern they found was that AI is well suited to winnowing rather than deciding. Used to cut the bottom 25 percent of candidates, the LLMs retained every applicant that human raters would have selected. Used as the sole decision-maker at the most selective ratios, they matched only about 41 to 62 percent of the candidates humans chose.

How Deeper Signals Approaches This

At Deeper Signals, this research directly informs how the platform is designed. Sola, the platform's AI assessment assistant, is built to support human decision-making rather than to replace it. Sola's guardrails specifically prevent it from making employment decisions, and its outputs are explainable and traceable back to specific, validated assessment constructs rather than functioning as an opaque black box.

This reflects the actual lesson of the mechanical-versus-clinical research. The value of mechanical combination comes from consistency in weighing already-valid data, not from removing human accountability altogether. Sola combines validated personality and values data consistently and transparently, while leaving the final hiring decision, and the responsibility that comes with it, with the people making it.

Frequently Asked Questions

Does this research mean human interviewers should be removed from hiring entirely?
No. The finding is specifically about combining already-collected data, not about eliminating human involvement in the hiring process. Structured interviews conducted by trained humans remain a valid, evidence-based data source that can be combined mechanically with other validated inputs.

Is AI in hiring always more accurate than human recruiters?
Not automatically. Mechanical combination outperforms holistic human judgment specifically when combining valid data through a validated formula. An AI system built on invalid, biased, or poorly designed inputs will not outperform a thoughtful human process, regardless of its algorithmic sophistication.

What is the difference between mechanical and clinical prediction?
Clinical prediction refers to an expert combining evidence about a candidate through holistic, subjective judgment. Mechanical prediction refers to combining the same evidence using a fixed, predetermined statistical formula. Kuncel et al. (2013) found mechanical combination more accurate on average.

Can an algorithm be legally accountable for a hiring decision?
No. Legal accountability for employment decisions remains with the organization and the humans who deploy and oversee the tool, regardless of how much of the process is automated. This is a key reason human oversight of AI hiring tools remains a legal and ethical necessity.

Should organizations trust AI hiring tools without human review?
No. Even well-validated algorithmic tools should be deployed with ongoing human oversight, regular fairness audits, and clear accountability for the final decision. The research supports using mechanical combination as one input into the process, not as an unsupervised decision-maker.

Last reviewed by Dr. Reece Akhtar — June 2026

References

Kuncel, N. R., Klieger, D. M., Connelly, B. S., & Ones, D. S. (2013). Mechanical versus clinical data combination in selection and admissions decisions: A meta-analysis. Journal of Applied Psychology, 98(6), 1060–1072.

Stockdale, K., Hickman, L., & Liu, S. (2026). Scoring employment interviews with large language models: Evaluation design components, validity investigations, and best practice recommendations. Journal of Applied Psychology. https://doi.org/10.1037/apl0001396

Subscribe
Subscribe to the Deeper Signals newsletter
Thank you! Your submission has been received!
Please fill all fields before submiting the form.
Curious to learn more?

Schedule a call with Deeper Signals to understand how our assessments and feedback tools help people gain a deep awareness of their talents and reach their full potential. Underpinned by science and technology, we build talented people, leaders and companies.

  • Scalable and engaging assessment solutions
  • Measurable and predictive talent insights
  • Powered by technology and science that drives results
Let's talk!
  • Scalable interventions for growth
  • Measurable data, insights and outcomes for high performance
  • Proven scientific expertise that links results to outcomes
Thank you!
Would you like to schedule a meeting?
Please fill all fields before submitting the form.