Can You Fake a Personality Test?
Yes, applicants can and do try to inflate their responses on personality assessments. The more important question is whether that inflation undermines the validity of personality testing enough to matter for talent decisions. The research evidence consistently shows that faking is real, that it occurs in selection contexts, and that its impact on predictive validity is more limited than critics typically assume. Understanding the evidence on both sides of this question is what responsible use of personality data requires.
Does Faking Actually Occur?
Birkeland, Manson, Kisamore, Brannick, and Smith (2006) conducted a meta-analysis specifically examining whether job applicants score differently on personality measures than non-applicant comparison groups, including students, incumbents, or respondents in research settings. Their findings confirmed that faking occurs in real selection contexts. Applicants scored meaningfully higher than comparison groups on most personality dimensions, with conscientiousness and emotional stability showing the largest differences. The effect was consistent across studies and instrument types.
This finding aligns with what any practitioner might intuitively expect: when candidates know that higher scores on desirable traits may help them secure a position, some will present themselves more favorably than they would under anonymous conditions. The research confirms this is a measurable and consistently observed phenomenon.
How Much Can Scores Be Inflated?
Viswesvaran and Ones (1999) conducted a meta-analysis of fakability estimates. They researched studies examining how much personality scores shift when respondents are instructed to fake good or respond as a desirable candidate would. Their analysis found that scores can be meaningfully inflated under deliberate faking conditions, with the largest effects on conscientiousness and emotional stability, the two dimensions most closely associated with socially desirable workplace behavior.
The magnitude of inflation is bounded, however. Respondents do not uniformly maximize every scale. Faking tends to be selective, reflecting implicit models of what a desirable candidate looks like in a particular context. Viswesvaran and Ones (1999) also found substantial variability in fakability across instruments and item formats. Some designs produce more faking-resistant profiles than others.
Does Faking Undermine Validity?
This is the critical question, and the research gives a more nuanced answer than either "faking destroys validity" or "faking doesn't matter at all."
Ones, Viswesvaran, and Reiss (1996) published one of the most influential papers in this debate, arguing that the impact of social desirability on the operational validity of personality tests has been systematically overstated. Using meta-analytic data, they showed that personality assessments retain meaningful validity in applicant samples at levels comparable to research settings, despite the greater opportunity for faking.
The explanation is not that faking fails to inflate scores. It does. The explanation is that inflation tends to be relatively uniform across candidates. Most applicants who want to appear conscientious shift their scores in the same direction, which preserves the rank ordering that validity depends on. A candidate who is genuinely conscientious and inflates their score remains above a candidate who is moderately conscientious and inflates by a similar amount. Rank order preservation is what predictive validity depends on, and uniform inflation disrupts it less than selective or idiosyncratic inflation would.
Morgeson, Campion, Dipboye, Hollenbeck, Murphy, and Schmitt (2007) challenged this argument, noting that self-report formats are vulnerable to motivated distortion and that the "red herring" position depends on faking being relatively uniform. That assumption may not hold when candidates have access to coaching or when a desirable response profile is widely known.
The most defensible position is that faking is real and measurable, that its impact on validity is smaller than intuition suggests in typical selection contexts, and that instrument design and context both shape how much of a practical problem it becomes.
What Factors Influence the Degree of Faking?
Several factors reliably affect how much faking occurs and how much it matters in practice.
Applicant motivation. Faking is higher when the stakes of the assessment are explicit and the desired profile is clear. High-stakes selection for competitive positions produces more faking than low-stakes developmental assessments. Birkeland et al. (2006) found that the magnitude of applicant-incumbent differences varied with the visibility of the assessment's selection purpose.
Instrument design. Forced-choice formats, in which respondents must choose between options matched for social desirability, are more resistant to faking than standard Likert-style formats, because they make it harder to systematically endorse all desirable options simultaneously. Response consistency indices, which flag statistically unlikely response patterns, can also detect signatures of deliberate distortion.
Trait visibility. Conscientiousness and emotional stability show the largest faking effects because their desirable ends are the most culturally legible. Almost everyone knows that organized, reliable, and calm people are valued at work. Traits with less obvious desirability hierarchies, such as openness to experience, show smaller faking effects.
Coaching and preparation. When detailed guidance on how to answer personality tests is available, whether from coaches, online resources, or organizational insiders, the uniformity assumption underlying the Ones et al. (1996) argument becomes more fragile. Differential access to coaching can produce non-uniform inflation that reduces rank order preservation and, by extension, reduces validity.
Is Faking Always a Bad Thing?
Before treating faking as a flaw to be eliminated, it is worth questioning the premise. Faking, more precisely called impression management, is something every person does and is something the workplace actively requires. Reading a situation, understanding what it calls for, and adjusting how you present yourself is not deviance, it is social competence. The candidate who tempers their bluntness in an interview, the manager who projects calm in a crisis, and the salesperson who mirrors a client are all managing impressions, and all are doing their jobs well. The line between distorting a personality test and demonstrating emotional intelligence is far thinner than the critics of self-report assume.
The research supports this reframing. Geiger, Bärwaldt, and Wilhelm (2021) found that the ability to fake successfully is a measurable individual difference that tracks a general factor related to perceiving and reading others, holding even after controlling for general mental ability. Pelt, van der Linden, and Born (2018) found that trait emotional intelligence positively predicts faking ability, with incremental validity beyond both cognitive ability and the Big Five.
The link to raw cognitive ability has also been studied. Schilling and colleagues (2021) found meta-analytically that personality scores correlate more strongly with cognitive ability in real selection settings than in low-stakes settings, because more able candidates are better at identifying and presenting the profile a role calls for. When someone successfully manages their impression on a well-designed assessment, they are demonstrating the same perceptiveness and social intelligence that organizations are trying to hire.
What Can Be Done About It?
The most defensible response to faking is not a single safeguard but a layered system, where multiple mechanisms each reduce the opportunity or incentive to distort and no single feature carries the full weight. The levers below are mutually reinforcing, and the practical impact of faking falls fastest when they are combined.
Forced-choice, desirability-matched design. Forced-choice item designs pair descriptors matched for social desirability, making it significantly harder to present an entirely idealized profile. Because both options are equally attractive, the simple strategy of endorsing every desirable statement no longer works.
Rapid-response and response-latency methods. Presenting items quickly and requiring fast responses constrains the time available for deliberate self-presentation. Meade, Pappalardo, Braddy, and Fleenor (2020) found that a rapid-response paradigm was significantly harder to fake than a traditional survey while retaining adequate reliability and validity.
Response-pattern and consistency analytics. Statistical checks identify respondents whose answers are internally incoherent, careless, or aberrant relative to genuine responding, flagging cases for review rather than automatically invalidating scores.
Norm-referenced scoring. Scoring candidates against a relevant norm group rather than in absolute terms preserves the rank ordering that predictive validity depends on. Because faking tends to be relatively uniform across candidates, normalizing scores against an applicant or working-population reference limits the degree to which uniform inflation distorts a person's relative standing.
Adaptive designs and test security. Adaptive item selection and controlled item exposure mean candidates do not all see the same fixed set of questions, which makes it far harder for coaching, leaked item banks, or shared answer keys to translate into a systematically inflated profile.
Restricting automated and proxy completion. Ensuring that the assessed person, rather than an AI agent, a third party, or an automated tool, is the one actually responding protects against a newer faking vector that bypasses the psychometric safeguards entirely.
Behaviorally framed item content. Item content that focuses on specific behavioral frequencies rather than evaluative trait labels ("how often do you complete tasks ahead of schedule?" rather than "are you organized?") is harder to respond to strategically without self-knowledge of actual behavior.
Context of use. Personality assessments used in developmental rather than selection contexts, where candidates understand that the data will be used for their own growth rather than to screen them out, consistently show lower faking effects and higher self-disclosure accuracy.
How Deeper Signals Approaches This
Deeper Signals’ assessments are designed with faking resistance as an explicit psychometric requirement, and Deeper Signals treats it as a layered system rather than a single safeguard.
Forced-choice, desirability-matched format. Our assessments use a forced-choice adjective-pair format in which both options in each pair are matched for social desirability. Respondents must choose between equally appealing or equally neutral descriptors, removing the straightforward "choose the most desirable option" strategy that makes Likert-based formats vulnerable.
Rapid-response paradigm. The diagnostic presents adjective pairs in rapid succession and asks for quick choices, constraining the time available for deliberate impression management. This mirrors the method validated by Meade, Pappalardo, Braddy, and Fleenor (2020).
Score normalization. Results are normalized against relevant norm groups rather than scored in absolute terms. Because faking inflation tends to be relatively uniform across candidates, normalization preserves the rank ordering that drives predictive validity and limits the degree to which any shared upward shift distorts a person's relative standing.
Algorithmic response checks. The platform includes algorithmic checks to detect careless, inconsistent, and aberrant responding, flagging patterns that are statistically unlikely under genuine self-report.
Restricting automated and proxy completion. Deeper Signals restricts the use of AI agents and other automated or third-party tools from taking assessments on a person's behalf, ensuring that the data reflects the actual individual being assessed and protecting test integrity against this emerging vector.
Adaptive design and test security. The assessment uses adaptive design elements that protect test security and limit item exposure, making it harder for coaching or shared item knowledge to produce a systematically inflated profile.
Frequently Asked Questions
Can candidates be coached to fake personality tests?
Coaching that teaches candidates to present a generically desirable profile can increase score inflation. However, uniform coaching across candidates preserves rank ordering, which limits its impact on validity. Coaching that teaches candidates to respond to specific instruments in ways that differ from their genuine profiles is harder to sustain under consistency checks and follow-up behavioral assessment.
Which personality dimensions are most fakeable?
Conscientiousness and emotional stability show the largest faking effects in meta-analytic research (Viswesvaran & Ones, 1999; Birkeland et al., 2006). These are the dimensions most closely associated with culturally legible workplace desirability. Openness to experience shows smaller faking effects because its desirable end is less universally apparent.
Does faking in personality tests hurt the organization or the candidate more?
Candidates who successfully present an inflated profile may be selected into roles that don't match their genuine behavioral tendencies, which typically surface as performance problems, dissatisfaction, or early attrition. The cost falls on both parties. This is one reason why personality data used for development, where accurate self-description serves the individual's interests, tends to show higher self-disclosure than data used for selection.
Are forced-choice personality tests always better than Likert formats?
Forced-choice formats reduce faking but introduce other psychometric trade-offs, including ipsative scoring properties that can complicate population-level comparisons. The best instrument design depends on the intended use case. For high-stakes selection, faking resistance is a priority. For developmental assessment, other properties, such as score reliability and feedback interpretability, may take precedence.
What is a response consistency index?
A response consistency index is a statistical measure of whether an individual's response pattern across a personality assessment is internally coherent. Patterns that are statistically improbable under genuine responding, for example, endorsing items that logically contradict each other at a rate far below chance, are flagged as potentially reflecting careless or deliberately distorted responding. Consistency indices do not prove faking but identify responses that warrant review.
Can an AI tool or another person take a personality test on a candidate's behalf?
This is a real and growing concern, since an automated agent or third party can bypass psychometric safeguards entirely. Deeper Signals restricts proxy completion and combines this with rapid-response timing and response-pattern analytics to protect data integrity.
Does normalizing scores actually help with faking?
Yes, indirectly but meaningfully. Predictive validity depends on the rank ordering of candidates, not on absolute scores. Because faking inflation tends to be relatively uniform across candidates, scoring people against a relevant norm group rather than in absolute terms preserves that rank ordering and limits the extent to which a shared upward shift changes who stands out. Normalization is not a standalone fix, but it is an important part of a layered approach.
Is any personality assessment completely fake-proof?
No, and any vendor who claims otherwise should be treated with caution. The right standard is defensibility, not perfection: a layered set of mitigations combined with validity evidence in real applicant samples. Deeper Signals is also expanding beyond self-report, with proctored assessment and 360 evaluations coming soon to add independent observer signals alongside self-report data.
Last reviewed by Dr. Reece Akhtar — June 2026
References
Ones, D. S., Viswesvaran, C., & Reiss, A. D. (1996). Role of social desirability in personality testing for personnel selection: The red herring. Journal of Applied Psychology, 81(6), 660–679. https://doi.org/10.1037/0021-9010.81.6.660
Viswesvaran, C., & Ones, D. S. (1999). Meta-analyses of fakability estimates: Implications for personality measurement. Educational and Psychological Measurement, 59(2), 197–210.
Morgeson, F. P., Campion, M. A., Dipboye, R. L., Hollenbeck, J. R., Murphy, K., & Schmitt, N. (2007). Reconsidering the use of personality tests in personnel selection contexts. Personnel Psychology, 60(4), 683–729.
Birkeland, S. A., Manson, T. M., Kisamore, J. L., Brannick, M. T., & Smith, M. A. (2006). A meta-analytic investigation of job applicant faking on personality measures. International Journal of Selection and Assessment, 14(4), 317–335. https://doi.org/10.1111/j.1468-2389.2006.00354.x
Geiger, M., Bärwaldt, R., & Wilhelm, O. (2021). The good, the bad, and the clever: Faking ability as a socio-emotional ability? Journal of Intelligence, 9(1), 13.
Meade, A. W., Pappalardo, G., Braddy, P. W., & Fleenor, J. W. (2020). Rapid response measurement: Development of a faking-resistant assessment method for personality. Organizational Research Methods, 23(1), 181–207. https://doi.org/10.1177/1094428118795295
Pelt, D. H. M., van der Linden, D., & Born, M. Ph. (2018). How emotional intelligence might get you the job: The relationship between trait emotional intelligence and faking on personality tests. Human Performance, 31(1), 33–54. https://doi.org/10.1080/08959285.2017.1407320
Schilling, M., Becker, N., Grabenhorst, M. M., & König, C. J. (2021). The relationship between cognitive ability and personality scores in selection situations: A meta-analysis. International Journal of Selection and Assessment, 29(1), 1–18. https://doi.org/10.1111/ijsa.12314


