All guides

Can Candidates Use AI to Cheat on Assessments, and Does It Matter?

Author
Dr. Reece Akhtar
CEO and Co-founder at Deeper Signals
Last reviewed
06/2026

Yes, candidates can use generative AI to attempt to manipulate assessment results. The research evidence confirms this is possible and worth taking seriously. What the same research also shows is that specific, well-established design choices substantially reduce how much impact this can have. Large-scale data on real job applicants points the same way, with little evidence of actual score inflation so far. The underlying question is therefore less about panic and more about which instruments were built with this risk in mind from the start. Understanding both sides of this evidence is what responsible use of AI-era assessment data requires.

What the Research Actually Found

Phillips and Robie (2024) built an experimental framework involving 655 business students and multiple large language models, using established integrity and conscientiousness measures. They specifically compared two item formats: single-stimulus questions, where a respondent rates one statement at a time, and phrase-based forced-choice questions, where a respondent chooses between paired statements matched for desirability.

Their findings were specific and directly useful for assessment design. ChatGPT models were able to produce favorable personality profiles that scored higher than the human student population, confirming that LLMs can manipulate personality assessment scores when prompted to do so. Just as importantly, phrase-based forced-choice measures proved meaningfully more resistant to this manipulation than single-stimulus questions. The format of the assessment, not simply the presence of AI, determined how much manipulation was actually possible.

But Is This Actually Happening in Real Hiring?

Capability is not the same as widespread use. Dabdoub and Pool (2026) tested that distinction on archival scores from over a million job applicants across three widely used pre-employment assessments, two cognitive and one personality, for the 27 months before and after ChatGPT's release. If cheating were widespread, scores should have climbed. They did not. The changes were near zero, inconsistent in direction, and explained almost no variance, and younger applicants showed no meaningful gains either.

The authors explain this lack of score inflation through how the tests are built. Their timed cognitive test and one-at-a-time personality items both make copying questions into an outside tool costly. The point is not that no one uses AI. It is that a capable tool being available has not, so far, produced the score inflation many feared, which is why the sensible response is monitoring and sound design rather than panic.

Not All Assessments Are Equally Exposed

Dr. Luke Treglown, Director of AI and Assessment R&D at Deeper Signals, has argued that cognitive ability assessments and personality assessments face genuinely different exposure to AI-assisted manipulation, and conflating the two leads to the wrong mitigation strategy.

Cognitive ability tests are more exposed to a specific kind of risk, since large language models are effective at finding correct answers to well-defined problems. Personality and soft-skill assessments face a different and more manageable risk. There is rarely an objectively correct answer to a personality item, so an AI asked to complete one without specific coaching toward a target profile tends to produce a generic or averaged result rather than anything strategically advantageous. The meaningful risk with personality assessments is a candidate directing AI toward a specific desirable profile, which is precisely the scenario forced-choice design measurably resists.

The exposure of cognitive ability tests is a genuinely different design problem, and it is addressed differently in practice. Rather than trying to write items an AI cannot answer, which is a difficult and often temporary fix given how quickly model capability improves, the more durable mitigation is controlling the conditions under which the test is taken. Restricting access to external tools during a proctored or time-constrained administration, using adaptive item selection so candidates do not all see the same fixed set of questions, and limiting how much any single item is exposed across administrations all reduce the practical opportunity for AI assistance, regardless of how capable the underlying model becomes.

Is This Cheating, or a Different Kind of Skill?

Dr. Treglown raises a genuinely useful reframe: if a candidate uses AI to help solve a complex problem within an assessment and produces an excellent solution, have they cheated, or have they demonstrated a real and increasingly relevant workplace skill, the ability to use available tools effectively to produce a good outcome?

This question matters for how assessments are designed going forward, but it does not change what responsible practice looks like today. For constructs where the goal is measuring a stable underlying trait, such as personality or values, an AI-generated response that does not reflect the candidate's own disposition is not useful data regardless of how the question is philosophically framed. The practical response is the same either way: build assessments that measure what they are designed to measure, using formats that make it difficult to substitute a generic or coached response for a genuine one.

What Can Be Done About It

The most defensible response is not a single safeguard but a layered system, where multiple mechanisms each reduce the opportunity for AI-assisted manipulation and no single feature carries the full weight.

Forced-choice, desirability-matched design. Phillips and Robie (2024) provide direct empirical evidence that this specific format measurably reduces how much LLM assistance can inflate a score, compared to single-stimulus formats.

Rapid-response formats. Presenting items quickly and requiring fast responses constrains the practical opportunity to consult an external tool mid-assessment, a benefit that predates the specific concern about generative AI but applies directly to it.

Clear candidate expectations, stated with a reason. Research on deterrence consistently finds that explicitly asking candidates not to use AI, and explaining why, meaningfully reduces misuse. People are more likely to comply with a request when they understand its purpose.

Restricting automated and proxy completion. Ensuring the assessed person, rather than an AI agent or third party, is the one actually responding closes off the most direct version of this risk.

Response consistency analytics. Statistical checks that flag internally incoherent or statistically improbable response patterns provide an additional layer of review without relying on any single indicator.

How Deeper Signals Approaches This

Deeper Signals' assessments were designed with faking resistance as an explicit psychometric requirement well before generative AI became a widely accessible tool, and that same layered design applies directly to this newer risk.

The Core Drivers Diagnostic and Core Values Diagnostic use a forced-choice, desirability-matched adjective-pair format, the exact design principle Phillips and Robie (2024) found measurably more resistant to LLM-assisted manipulation than single-stimulus alternatives. The assessment also presents items in a rapid-response format that constrains the time available for consulting an outside tool, and results are normalized against relevant norm groups rather than scored in absolute terms, which limits how much any shared upward shift can distort a candidate's relative standing.

The Core Reasoning Suite, the platform's cognitive ability assessment, addresses the different exposure that cognitive tests specifically face. It uses adaptive item selection, so candidates do not all see the same fixed set of questions, which limits how effectively a shared or leaked answer set could be exploited across candidates. Time-constrained administration and restrictions on external tool access during testing further reduce the practical opportunity to consult AI mid-assessment, addressing the exposure at the level of testing conditions rather than relying solely on item design to outpace model capability.

Deeper Signals restricts the use of AI agents and other automated or third-party tools from completing assessments on a candidate's behalf, and combines this with algorithmic checks that flag statistically improbable response patterns for review. The platform also uses adaptive design elements that limit item exposure, making it harder for a shared or leaked item set to be exploited at scale. Together, they reflect the standard the evidence actually supports: a validated format, clear expectations, and ongoing monitoring, applied in combination.

Frequently Asked Questions

Can AI actually complete a personality test convincingly?
Yes, under some conditions. Phillips and Robie (2024) found that ChatGPT models could produce personality profiles that scored higher than human respondents when prompted to do so, particularly on single-stimulus item formats.

Have AI tools actually inflated assessment scores in real hiring?
So far, the evidence points to no. Dabdoub and Pool (2026) tracked over a million real applicants and saw no practically meaningful score inflation after ChatGPT's release.

Does using a forced-choice format make an assessment immune to AI manipulation?
No format eliminates the risk entirely, but forced-choice, desirability-matched design measurably reduces it. Phillips and Robie (2024) found phrase-based forced-choice measures significantly more resistant to LLM-driven score inflation than single-stimulus questions.

Are cognitive ability tests or personality tests more vulnerable to AI cheating?
They are exposed in different ways and mitigated differently. Cognitive ability tests are exposed because AI is effective at finding correct answers to well-defined problems, and this is best addressed through testing conditions, such as controlled administration and adaptive item selection. Personality tests are exposed differently, since a candidate can direct AI to fake a specific, desirable profile, which a well-designed forced-choice format meaningfully limits.

Should organizations rely on webcam proctoring to prevent AI-assisted cheating?
Proctoring can reduce some misuse, but it is one tool among several rather than a standalone solution, and it carries its own trade-offs for candidate experience. A validated assessment format that is inherently more resistant to manipulation reduces reliance on surveillance in the first place.

Is any personality assessment completely immune to AI-assisted manipulation?
No, and any provider claiming otherwise should be treated with caution. The appropriate standard is defensibility, not perfection: a layered set of design choices and monitoring practices, combined with ongoing validity evidence in real applicant samples.

Last reviewed by Dr. Reece Akhtar — June 2026

References

Dabdoub, A., & Pool, R. (2026). Testing after ChatGPT: Aggregate change in applicant assessment scores. International Journal of Selection and Assessment, 34, e70077. https://doi.org/10.1111/ijsa.70077

Phillips, J. J., & Robie, C. (2024). Hacking the perfect score on high-stakes personality assessments with generative AI. Personality and Individual Differences, 231, 112840. https://doi.org/10.1016/j.paid.2024.112840

Subscribe
Subscribe to the Deeper Signals newsletter
Thank you! Your submission has been received!
Please fill all fields before submiting the form.
Curious to learn more?

Schedule a call with Deeper Signals to understand how our assessments and feedback tools help people gain a deep awareness of their talents and reach their full potential. Underpinned by science and technology, we build talented people, leaders and companies.

  • Scalable and engaging assessment solutions
  • Measurable and predictive talent insights
  • Powered by technology and science that drives results
Let's talk!
  • Scalable interventions for growth
  • Measurable data, insights and outcomes for high performance
  • Proven scientific expertise that links results to outcomes
Thank you!
Would you like to schedule a meeting?
Please fill all fields before submitting the form.