All guides

What Is Reliability in Psychometric Testing?

Author
Dr. Reece Akhtar
CEO and Co-founder at Deeper Signals
Last reviewed
06/2026

Reliability in psychometric assessments refers to the consistency of measurement, the degree to which an assessment produces stable, reproducible scores across items, occasions, and raters. An assessment is reliable when it produces the same result for the same person under equivalent conditions. Reliability is a necessary condition for validity. An instrument that measures inconsistently cannot produce consistent predictions of job performance or any other outcome. Reliability is one of the two foundational technical requirements, alongside validity, that any psychometric instrument must satisfy before it is suitable for professional use.

Why Reliability Matters

The practical consequence of unreliable measurement is noise, random variation in scores that is attributable to the measurement process rather than to genuine differences in the construct being assessed. When an assessment is unreliable, a candidate's score reflects some combination of their true standing on the measured trait and random measurement error. That error component reduces predictive validity: a score contaminated by measurement noise predicts job performance less accurately than a score that precisely reflects the candidate's true trait level.

Schmidt and Hunter (1998) demonstrated this relationship formally through the statistical concept of attenuation: measurement error in either the predictor (the assessment) or the criterion (the performance measure) systematically reduces the observed correlation between them. Meta-analytic syntheses of selection method validity routinely apply reliability corrections to raw validity estimates to produce estimates that better reflect the true relationship between assessment scores and performance, free of the attenuating effects of measurement imprecision. Understanding reliability is therefore not merely a technical exercise. It is a prerequisite for interpreting the validity evidence that informs selection decisions.

Types of Reliability

Reliability is not a single property. Several distinct forms capture different aspects of measurement consistency, and each is relevant for different applications and interpretive questions (Kline, 1994).

Internal consistency reliability measures the degree to which items within a scale intercorrelate, the extent to which the items are consistently measuring the same construct. It is the most widely reported reliability statistic in psychometric research and is most commonly expressed using Cronbach's alpha, a coefficient ranging from 0 to 1. Internal consistency reliability reflects whether the items in a scale hang together as a coherent set.

Test-retest reliability measures the stability of scores over time. An individual completing the same assessment at two different points in time should produce similar scores if the assessment measures a stable trait and the trait itself has not genuinely changed. Test-retest reliability is expressed as a correlation between Time 1 and Time 2 scores, and it is particularly important for personality assessments and cognitive ability tests, which are designed to measure constructs that are relatively stable in adults. An instrument that produces substantially different scores from the same person across two administrations separated by weeks or months is capturing measurement noise rather than the stable trait it claims to measure.

Inter-rater reliability measures the degree of agreement between two or more independent raters scoring the same responses. It is most relevant for assessments that involve human judgment in scoring, such as structured interview ratings, assessment center exercises, and behavioral observation measures. Inter-rater reliability is expressed either as a correlation between rater scores or as a kappa coefficient that adjusts for chance agreement. When inter-rater reliability is low, scores reflect the individual rater's idiosyncratic judgments as much as the candidate's actual performance.

Understanding Cronbach's Alpha

Cronbach's alpha is the reliability statistic most frequently encountered in psychometric technical manuals, and it is also the most frequently misinterpreted. Cortina (1993), in an influential examination of coefficient alpha and its applications in Journal of Applied Psychology, identified several specific ways practitioners misuse this statistic.

Alpha measures internal consistency, the average intercorrelation among items in a scale, corrected for scale length. It does not measure unidimensionality. A scale can produce a high alpha while measuring two or more distinct constructs simultaneously, if those constructs correlate with each other. Alpha also does not measure test-retest stability. A scale with high internal consistency can produce very different scores from the same person on different occasions, which means high alpha does not confirm that the assessment is measuring a stable trait.

Alpha is also sensitive to scale length in ways that can mislead. Adding more items to a scale increases alpha mechanically, even if those items are redundant or only weakly related to the construct of interest. A scale with 30 highly redundant items will produce a higher alpha than a scale with 10 carefully selected items that cover the construct's full range. However, the longer scale is not necessarily more valid or more useful.

The professional standard for internal consistency reliability is a Cronbach's alpha of .70 or above for applied use (Nunnally, 1978; Kline, 1994). Values above .80 are considered strong. Values above .90 may indicate excessive item redundancy and should be examined alongside dimensionality analysis rather than accepted uncritically. It is also worth noting that alpha assumes all items contribute equally to the construct (an assumption known as tau-equivalence) that real scales rarely meet, which is why modern psychometric practice increasingly reports McDonald's omega alongside or instead of alpha as a more accurate estimate of internal consistency.

Reliability and Validity: The Relationship

Reliability and validity are related but distinct properties, and understanding how they relate is essential for evaluating assessment quality. Reliability is a ceiling on validity: the maximum correlation an assessment can achieve with an external criterion is bounded by its own reliability. Specifically, the maximum observed correlation between two measures equals the square root of the product of their two reliabilities, meaning that unreliable measures produce attenuated validity estimates even when the true relationship is strong.

This relationship has a practical implication. An assessment with a reliability of .70, correlated with a perfectly reliable criterion (reliability of 1.0), cannot exceed an observed validity coefficient of about .84, because the square root of (.70 × 1.0) is .84. Real performance criteria are themselves imperfectly measured, so operational validity coefficients are always lower than this theoretical ceiling. Improving reliability directly raises the upper bound on validity that an instrument can achieve.

Reliability is necessary but not sufficient for validity. A perfectly reliable assessment could consistently measure something entirely irrelevant to the outcomes of interest, producing stable, reproducible scores that predict nothing. The Standards for Educational and Psychological Testing (AERA, APA, & NCME, 2014) require that assessments used in professional settings document both reliability and validity evidence, because neither alone constitutes adequate evidence of fitness for use. For a buyer, the practical takeaway is to ask for both, and to be wary of any vendor that presents one without the other.

Reliability in Score Units: The Standard Error of Measurement

Reliability coefficients are abstract. The standard error of measurement (SEM) translates reliability into the score units a practitioner actually works with, which makes it the single most useful reliability concept for interpreting an individual result. The SEM expresses how much an individual's observed score would be expected to fluctuate across repeated administrations due to measurement error alone, and it is used to build a confidence band around a score. A score of 65 with an SEM that produces a band of plus or minus 5 points should be read as a range, not a precise point.

The practical consequence is that small differences between scores should not be over-interpreted. If two candidates score 62 and 65 on the same scale and the confidence bands overlap substantially, the difference is within measurement noise and should not drive a decision on its own. SEM is directly derived from reliability: the more reliable the scale, the narrower the confidence band, and the more confidently a practitioner can treat a score difference as real. A technical manual that reports SEM, or confidence intervals, alongside reliability coefficients is demonstrating a mature approach to score interpretation.

What Practitioners Should Look For

When evaluating a psychometric assessment for organizational use, practitioners should request reliability evidence across all three relevant types.

A technical manual for a professionally sound instrument will report internal consistency coefficients (Cronbach's alpha, and ideally McDonald's omega) for each scale, typically above .70. It will report test-retest reliability for personality and cognitive measures, with values above .70 across intervals of at least four to eight weeks, indicating adequate stability. For assessments involving human scoring, it will report inter-rater reliability statistics. It will report the standard error of measurement or confidence intervals so scores can be interpreted as ranges. And it will document how these reliability figures were obtained, including the sample size, the time interval for test-retest data, and the scoring conditions, rather than simply reporting a number without context.

As a quick reference, adequate benchmarks for applied use are: internal consistency at or above .70 (above .80 is strong); test-retest at or above .70 across a four to eight week interval; and, for human-scored measures, inter-rater agreement reported with a chance-corrected statistic such as kappa.

How Deeper Signals Approaches Reliability

The Core Drivers Diagnostic reports internal consistency (Cronbach's alpha) coefficients ranging from .68 to .82 across its six personality scales, with the large majority at or above the .70 benchmark and most in the strong range. Where a scale sits marginally below .70, this reflects a deliberate trade-off: the scales are built for broad construct coverage rather than for maximizing alpha through redundant items, and the modest internal-consistency figure is backed by strong stability over time. Test-retest reliability across a four-week interval falls in the range of r = .68 to .77, consistent with British Psychological Society guidance for personality measures and with the stability expected from a Five Factor Model instrument measuring traits intended to be relatively stable in adults.

The Core Drivers Diagnostic does not inflate alpha by padding scales with redundant items. The assessment was developed using genetic algorithms to select 90 items from an initial pool of 300, with item selection optimized for convergent validity and construct coverage rather than alpha maximization. This reflects Cortina's (1993) insight that high alpha alone is not a quality target, and it is the reason a scale can show broad coverage and strong test-retest stability without an artificially high alpha. Scores are reported with interpretive context rather than as bare point estimates, so that small, within-error differences are not over-read. This methodology is part of a consistent design philosophy across Deeper Signals assessments, alongside the faking-resistant response format and the large global norm base described in our companion articles.

Frequently Asked Questions

What is a good Cronbach's alpha for a personality assessment?

Values of .70 and above are generally considered adequate for applied professional use, and values of .80 and above are considered strong (Kline, 1994). Values above .90 may indicate item redundancy and should be examined alongside evidence about the scale's construct coverage and dimensionality.

Does high reliability mean an assessment is valid?

High reliability is a necessary but not sufficient condition for validity. A reliable assessment consistently measures something, but that something may or may not predict the outcomes of interest. Both reliability and validity evidence must be present in a technical manual before an assessment is considered fit for professional use.

Why does test-retest reliability matter for personality assessments?

Personality traits are understood to be relatively stable in adults. An instrument that produces substantially different personality profiles for the same person across two administrations separated by weeks is not measuring a stable trait. It is reflecting measurement noise rather than genuine individual differences. Test-retest reliability confirms that the assessment captures stable characteristics rather than situational fluctuations.

What is the difference between reliability and consistency?

In psychometric terminology, reliability and consistency are closely related. Internal consistency specifically refers to item intercorrelations within a scale. Reliability is the broader concept encompassing internal consistency, stability over time, and agreement between raters. Both refer to the reproducibility of measurement, but they capture different sources of potential inconsistency.

Can an assessment have high test-retest reliability but low internal consistency?

Yes. An assessment measuring a stable construct with items that tap different facets of that construct can show high stability over time while items correlate only moderately with each other, producing adequate test-retest reliability but modest alpha. This is not necessarily a problem; it may reflect genuine breadth of construct coverage. This is one reason why Cortina (1993) emphasizes that alpha should not be treated as the sole indicator of reliability quality.

Last reviewed by Dr. Reece Akhtar — June 2026

References

Kline, P. (1994). The Handbook of Psychological Testing. London: Routledge.

Nunnally, J. C. (1978). Psychometric Theory (2nd ed.). New York, NY: McGraw-Hill.

Cortina, J. M. (1993). What is coefficient alpha? An examination of theory and applications. Journal of Applied Psychology, 78(1), 98–104. https://doi.org/10.1037/0021-9010.78.1.98

Schmidt, F. L., & Hunter, J. E. (1998). The validity and utility of selection methods in personnel psychology: Practical and theoretical implications of 85 years of research findings. Psychological Bulletin, 124(2), 262–274. https://doi.org/10.1037/0033-2909.124.2.262

American Educational Research Association, American Psychological Association, & National Council on Measurement in Education. (2014). Standards for Educational and Psychological Testing. Washington, DC: American Educational Research Association.

Subscribe
Subscribe to the Deeper Signals newsletter
Thank you! Your submission has been received!
Please fill all fields before submiting the form.
Curious to learn more?

Schedule a call with Deeper Signals to understand how our assessments and feedback tools help people gain a deep awareness of their talents and reach their full potential. Underpinned by science and technology, we build talented people, leaders and companies.

  • Scalable and engaging assessment solutions
  • Measurable and predictive talent insights
  • Powered by technology and science that drives results
Let's talk!
  • Scalable interventions for growth
  • Measurable data, insights and outcomes for high performance
  • Proven scientific expertise that links results to outcomes
Thank you!
Would you like to schedule a meeting?
Please fill all fields before submitting the form.