All guides

What Is an AI Assessment Tool?

Author
Dr. Reece Akhtar
CEO and Co-founder at Deeper Signals
Last reviewed
06/2026

An AI assessment tool is any instrument that uses machine learning or other computational methods to score, predict, or interpret information about a candidate or employee. This is a broader category than most people realize, and the tools within it vary enormously in how rigorously they have been validated. Some AI assessment tools apply machine learning directly to raw behavioral data, such as scoring personality from a video interview. Others apply AI to already-validated psychometric data, using it to generate interpretation or coaching guidance rather than to produce the underlying score itself. These are fundamentally different categories of tool, and evaluating them requires understanding which one is actually being used.

The Two Categories of AI Assessment Tool

Machine learning-scored assessments use AI to generate the score itself, directly from raw behavioral data such as video, audio, or written text. An automated video interview tool that analyzes facial expressions, word choice, and vocal tone to produce a personality score is in this category. The AI is doing the actual measurement.

AI-assisted interpretation of validated assessments uses AI to explain, contextualize, or apply a score that was generated through an established, validated psychometric method, typically self-report or structured behavioral rating. The AI is not measuring the underlying trait. It is helping translate an already-produced score into something more usable.

This distinction matters because the evidence bar for each category is different. A tool that generates the score itself carries the full weight of proving that its measurement process is valid. A tool that interprets an existing validated score inherits the validity of that underlying instrument and adds a separate, narrower question: does the interpretation layer itself add value without distorting the original data?

What the Research Shows About Machine Learning-Scored Assessments

Hickman, Bosch, Ng, Saef, Tay, and Woo (2022) conducted the most rigorous published investigation of automated video interviews that use machine learning to score Big Five personality traits from verbal, paraverbal, and nonverbal behavior. Their study, which won the Society for Industrial and Organizational Psychology's 2024 Jeanneret Award for excellence in individual assessment research, developed and tested machine learning models across three separate samples of video interviews, totaling more than 1,000 participants, and examined test-retest reliability in a fourth sample.

Their most important and frequently overlooked finding was that these automated personality assessments showed meaningfully stronger validity evidence when the underlying machine learning models were trained using interviewer ratings as the target, rather than candidates' own self-reported personality scores. This has a direct practical implication. An AI tool that is simply trained to predict what a candidate says about themselves inherits the accuracy limitations of self-report, whereas a tool trained on expert observer judgments has the potential to capture something closer to how the candidate is actually experienced by others.

This finding also illustrates a broader point about machine learning-scored assessments generally. Their validity depends heavily on what the model was trained to predict, not simply on how sophisticated the underlying algorithm is. A technically advanced model trained on a weak or poorly defined target will produce a weak or poorly defined score, regardless of how much data it was trained on.

A second, independent line of research reaches a similar conclusion. Speer, Delacruz, Chawota, Perrotta, and Rudolph (2026) studied large language models scoring open-ended personality narratives, where people write about themselves rather than answer fixed items. Zero-shot, generic AI scoring performed worst, while fine-tuned models performed best, and every method still showed at least some validity.

What mattered most was not the model alone but the rigor built around it. Validity depended heavily on how the writing prompts were designed, with prompts crafted to draw out trait-relevant detail clearly outperforming generic ones. When the models were trained to predict direct evaluations of the responses, their scores correlated strongly with the target scores, averaging about .83.

Why Construct Clarity Still Matters

Dr. Luke Treglown, Director of AI and Assessment R&D at Deeper Signals, makes a related point about how AI should be used in building assessments in the first place. He argues that AI does not replace the psychologist's role in defining what a construct actually is. If anything, it raises the importance of that role. A psychologist who clearly defines a construct and its sub-components gives an AI system a much stronger foundation to work from, whether that AI is generating candidate items, extracting signal from video and text, or scoring a construct directly.

This means the quality of an AI assessment tool cannot be evaluated by its technical sophistication alone. It has to be evaluated by whether the underlying construct was clearly and rigorously defined before any machine learning was applied to it.

Evaluating Any AI Assessment Tool

Ask what the AI is actually scoring, and what it was trained to predict. A machine learning-scored assessment trained on weak or ambiguous targets will produce a score that is only as good as that target, regardless of algorithmic sophistication.

Request the same evidence you would require of any assessment. Criterion validity, reliability, normative data, and adverse impact analyses are not optional simply because a tool uses AI. If anything, machine learning-scored tools require additional scrutiny because their scoring logic is often less transparent than a traditional psychometric instrument.

Check whether the tool has been tested for generalizability across contexts. Hickman et al. (2022) specifically investigated whether automated video interview scores held up across different interview questions and settings, not only within the specific sample the model was trained on. A tool validated in one narrow context may not generalize to your specific population or use case.

Distinguish AI that measures from AI that interprets. These require different evidence. A tool generating a score directly from behavioral data needs to prove its measurement process is valid. A tool interpreting an already-validated score needs to prove that its interpretation adds value without distorting the underlying data.

Ask whether the tool gives the same person the same score every time. Boyce, Hickman, and Boyce (2026) flag a problem specific to AI scoring: a candidate can get a different score just because a different AI model or setting was used, or it was run again. A traditional questionnaire does not vary in this way. A credible vendor should be able to show that its scores are stable and repeatable, so a result does not come down to chance.

How Deeper Signals Approaches This

At Deeper Signals, the platform deliberately separates measurement from interpretation. The underlying personality and values scores are generated through validated, self-report psychometric instruments, the Core Drivers Diagnostic and Core Values Diagnostic, not through machine learning applied directly to video, audio, or unstructured behavioral data. This means the measurement itself carries the documented reliability, validity, and adverse impact evidence that these instruments have been tested against.

Sola, the platform's AI assessment assistant, sits entirely in the second category. It does not generate assessment scores. It interprets and applies scores that have already been produced through validated psychometric methods, turning them into personalized, real-time guidance. This design reflects the distinction this post makes directly. Sola's role is to make already-valid data more useful, not to introduce a new, less transparent measurement process into the pipeline.

Frequently Asked Questions

Are all AI assessment tools equally valid?
No. Validity depends heavily on what the tool was trained to measure, how it was tested, and whether it generalizes across different contexts and populations. Hickman et al. (2022) found meaningful validity differences even within a single category of tool, automated video interviews, depending on how the underlying models were trained.

Is a machine learning-scored assessment inherently less trustworthy than a traditional one?
Not inherently, but it requires more scrutiny in specific areas, particularly transparency about what the model was trained to predict and evidence that its scores generalize across different contexts. A rigorously validated machine learning-scored tool can meet a high evidence bar, but that evidence needs to be requested and reviewed directly.

What should I ask a vendor about their AI assessment tool?
Ask specifically what the underlying model was trained to predict, what data it was trained on, how its scores compare to established psychometric benchmarks, and whether its validity and reliability have been tested across different samples and contexts, not only in the original development sample.

Can AI-generated interview scores be biased?
Yes. Machine learning models trained on behavioral data can learn and reproduce biases present in that training data, including biases related to accent, appearance, or communication style that have nothing to do with job-relevant capability. Adverse impact testing is essential for any AI tool that scores candidates directly from behavioral data.

Does using AI in an assessment tool always mean lower quality?
No. AI can meaningfully improve assessment tools when applied thoughtfully, whether in generating diverse item content, improving translation and localization, or making validated assessment data more accessible through interpretation. The quality of the outcome depends on how carefully the underlying construct was defined and how rigorously the tool was validated, not on whether AI was involved at all.

Last reviewed by Dr. Reece Akhtar — June 2026

References

Boyce, A. S., Hickman, L., & Boyce, C. E. (2026). The future of selection enabled by artificial intelligence. In N. Schmitt & A. M. Ryan (Eds.), The Oxford handbook of personnel assessment and selection (2nd ed.). Oxford University Press. https://doi.org/10.1093/9780197809013.003.0018

Hickman, L., Bosch, N., Ng, V., Saef, R., Tay, L., & Woo, S. E. (2022). Automated video interview personality assessments: Reliability, validity, and generalizability investigations. Journal of Applied Psychology, 107(8), 1323–1351. https://doi.org/10.1037/apl0000695

Speer, A. B., Delacruz, A. Y., Chawota, T. A., Perrotta, J., & Rudolph, C. W. (2026). Unpacking the validity of open-ended personality assessments using fine-tuned large language models. Organizational Research Methods. Advance online publication. https://doi.org/10.1177/10944281251413746

Subscribe
Subscribe to the Deeper Signals newsletter
Thank you! Your submission has been received!
Please fill all fields before submiting the form.
Curious to learn more?

Schedule a call with Deeper Signals to understand how our assessments and feedback tools help people gain a deep awareness of their talents and reach their full potential. Underpinned by science and technology, we build talented people, leaders and companies.

  • Scalable and engaging assessment solutions
  • Measurable and predictive talent insights
  • Powered by technology and science that drives results
Let's talk!
  • Scalable interventions for growth
  • Measurable data, insights and outcomes for high performance
  • Proven scientific expertise that links results to outcomes
Thank you!
Would you like to schedule a meeting?
Please fill all fields before submitting the form.