Predictive validity: the figures, and what sits behind them

Assessment is full of impressive numbers that rarely say what they were measured on. Here is the whole of NeuroFrame's psychometrics: every figure explained, every sample named, and an honest note about what these numbers do not prove.

What predictive validity is, and how it is measured

Predictive validity answers a plain question: how strongly does an assessment result, obtained before hiring, relate to how the person actually performs afterwards.

It is expressed as a correlation coefficient, r, running from 0 to 1. Zero means there is no relationship at all — the instrument tells you nothing a coin toss would not. One means perfect prediction, which does not occur anywhere in occupational psychology.

The second figure is R², the square of r. It answers a different question: what share of the differences in work performance the instrument accounts for. If r = 0.31, then R² = 0.10 — roughly a tenth of the variation is explained, and the remaining nine tenths belong to everything else: the brief, the manager, the team, the market, health, luck.

Two practical consequences follow. First, squaring a coefficient shrinks it considerably, so r and R² must never be confused with one another or set side by side. Second, in the meta-analytic table below, even the strongest selection methods account for less than a third of the differences in performance. That is not a failure of psychometrics but an accurate picture: how well someone works is never determined by that person alone.

One further point is worth holding on to. The same instrument yields different values of r depending on what was taken as the criterion — a supervisor's rating, a sales figure, retention at twelve months — and on the sample it was measured in. A validity coefficient is a statement about a particular study, not a permanent property of the method.

NeuroFrame's psychometric figures, and what each rests on

Psychometrics is the science of how accurately and how stably an instrument measures what it claims to measure. It has an inconvenient feature: each figure answers its own narrow question, and none of them on its own establishes that an instrument is sound. What follows is therefore the complete set, with the data behind each one.

On criterion validity specifically. The cognitive scales were checked against recognised tests — Raven's Progressive Matrices and Saville instruments — and against measures of attention, memory, task switching, information processing and decision-making. The personality scales were validated through established psychometric models: openness to experience, risk appetite, persistence, agreeableness and conscientiousness.

What this list does not contain, and will not: the exact weightings of the behavioural indicators, the statistical model that converts behaviour into scores, or the workings of the machine-learning component. The scoring key stays closed — not out of secrecy, but because the assessment reads behaviour that the person is not consciously putting forward for judgement. Publish the key and the measure becomes a test that can be prepared for, at which point the normative base loses its value for everyone assessed afterwards. Keeping scoring keys secure is a requirement of the professional standards for assessment (AERA/APA/NCME; the ITC Guidelines on Security of Tests). The full evidential logic — what is measured and through which behaviour it is captured — is disclosed to a client's own methodologists under a confidentiality agreement.

The figures, explained, with the data behind them

  • Comparison sample — 14,850. Executives and entrepreneurs from 500+ companies, 21 industries and 21 functions. This is the base against which an individual result is calculated: every score is normative, describing a position within a comparable group rather than an absolute quantity.
  • Validation sample — 3,000+ employees. A separate exercise: assessment results were set against real KPIs and managers' ratings. This is the data on which predictive value was tested under working conditions.
  • CFI ≈ 0.96 — construct validity. Confirmatory factor analysis supports a two-domain model, cognition and personality, as a good description of the observed data. The domains themselves were identified by exploratory analysis beforehand, so the structure was found in the data and then re-tested rather than imposed on it.
  • Cronbach's α = 0.69–0.77 — internal consistency. The indicators within each scale measure the same quality; the scale does not fall apart into unrelated pieces. It is a range rather than a single figure because there are several scales, each with its own value.
  • Test–retest reliability > 0.83. Results reproduce on a second sitting, which means the method captures stable characteristics rather than the mood of a particular Tuesday.
  • AUC ≈ 0.77 — the link to outcomes on our own data. How reliably the model separates groups with differing levels of work performance, measured on the NeuroFrame validation sample of 3,000+ employees. 0.5 is guesswork; 1.0 is faultless separation.
  • Machine-learning model accuracy — 72–89%. The spread exists because there are several models addressing different tasks. A single averaged figure would look tidier and describe the position less accurately.

Why our figures do not belong in the table below

What follows is a table of meta-analytic validity coefficients for common selection methods. Before reading it, one caveat — and we are putting it here rather than in small print underneath.

NeuroFrame's AUC ≈ 0.77 and the figures in the table were obtained on different samples and against different criteria. Ours describes how well the model performs on NeuroFrame's own data. The meta-analytic coefficients describe the relationship between a method and externally measured job performance, averaged across dozens of independent studies over decades. These are neighbouring quantities, not equivalent ones — AUC and the coefficient r answer different questions on different scales — and arithmetic of the form "our figure against their coefficient" does not hold here.

What the table does show properly is the order of magnitude across the industry. The strongest of the classical methods account for something like a quarter to a third of the differences in work performance. Personality questionnaires account for between one and ten per cent. Job tenure and years of education account for a few per cent, which is to say almost nothing. And the typologies best known to the general public, MBTI and DiSC, publish no predictive validity data at all — which is itself informative.

The table shows something else too: even settled figures get revised. The 2022 re-analysis by Sackett and colleagues applied a stricter correction for range restriction and brought the classical Schmidt–Hunter coefficients down by roughly a third. A validity coefficient is an estimate that depends on how it was calculated, not a constant.

Our own position is the one that can be checked: the method has been validated against recognised tests, against real KPIs and against managers' ratings. Setting that directly against someone else's meta-analysis is a comparison we are not going to make.

What the method does not do

An honest account of the limits says more about an instrument than another decimal place does.

The ranges are a reference point, not a verdict: a gap on one or two parameters is something to explore at interview, not grounds for rejection.

Where the method stops

  • The empirical base for game-based assessment is still accumulating. Validity figures depend on the design of the individual task, and each task is validated separately.
  • Interpretation is always normative: a result describes a position within a comparable group, not a value on an absolute scale.
  • The assessment does not replace interviews, verification of experience, or managerial judgement. It provides grounds for a decision; it is not the decision.
  • Unusual roles and distinctive corporate cultures call for calibration against the client's own high performers.
  • The method improves the odds of a sound decision. It does not guarantee that a particular person will succeed in a particular role — and no selection instrument does.

Predictive validity of common selection methods, according to two meta-analyses

Selection methodr — Schmidt & Hunter, 1998r — Sackett et al., 2022R² — share of variance explained
Cognitive ability tests (GMA)0.510.310.26 · 0.10
Structured interview0.510.26
Work samples0.540.29
Personality questionnaires, conscientiousness0.310.190.10 · 0.04
Job tenure0.180.03
Years of education0.100.01
MBTI, DiSCnot published

r is the validity coefficient; R² is its square — the share of differences in work performance that the method accounts for. Where the R² column carries two values separated by "·", they correspond to the two r columns in the same order. A dash means the source gives no separate estimate. The Sackett re-analysis applies a stricter correction for range restriction, which is why its values are systematically lower. Schmidt, F. L., & Hunter, J. E. (1998). The validity and utility of selection methods in personnel psychology. Psychological Bulletin, 124(2), 262–274. DOI: 10.1037/0033-2909.124.2.262 Sackett, P. R., Zhang, C., Berry, C. M., & Lievens, F. (2022). Revisiting meta-analytic estimates of validity in personnel selection. Journal of Applied Psychology, 107(11), 2040–2068. DOI: 10.1037/apl0000994 NeuroFrame's own figures are deliberately absent from this table: they were obtained on different samples and against different criteria, and are not open to direct comparison.

Common questions

How accurate is it?

As accurate as assessment of people ever gets. On a validation sample of 3,000+ employees the model reaches AUC ≈ 0.77: it separates groups with differing levels of work performance well clear of the 0.5 that marks guesswork. Machine-learning model accuracy is 72–89% depending on the model. Test–retest reliability is above 0.83, so results reproduce on a second sitting. That said, no assessment guarantees that a given person will succeed in a given role. It improves the odds of a sound decision, and nothing more.

What validity does the method have, and how was it established?

Construct validity — CFI ≈ 0.96: confirmatory factor analysis supports a two-domain model, cognition and personality. Internal consistency — Cronbach's α of 0.69–0.77 across the scales. Test–retest reliability — above 0.83. Criterion validity was established by correlating the cognitive scales with Raven's Progressive Matrices and Saville instruments, and with measures of attention, memory, task switching, information processing and decision-making; the personality scales were validated through established psychometric models. Predictive value was tested against real KPIs and managers' ratings on a sample of 3,000+ people. The normative comparison base comprises more than 10,000 executives and entrepreneurs from 500+ companies, 21 industries and 21 functions.

Is there an evidence base?

Yes, and it works on two levels. The first is the scientific grounding of the method: the Big Five, Gray's Reinforcement Sensitivity Theory, Eysenck's PEN model, Cybernetic Big Five Theory, Evidence-Centered Design and stealth assessment (Shute, 2011). The second is NeuroFrame's own psychometric data, set out on this page with the samples named. The full evidential logic — what is measured and through which behaviour it is captured — is disclosed to a client's own methodologists under a confidentiality agreement. Only the operational key stays closed: the indicator weightings, the model that converts behaviour into scores, and the machine-learning component. That is a condition of preserving validity rather than corporate reticence; keeping scoring keys secure is a requirement of the professional standards for assessment.

Are you more accurate than the classical tests?

Not a claim we will make in that form. Our own figures were obtained on our own data, whereas the meta-analytic coefficients for other instruments come from different samples and different criteria. Setting them directly against one another is not sound, and an informed reader would rightly challenge it. What can honestly be said is this: many popular methods publish no predictive validity data at all, and personality questionnaires account for only a small share of the differences in performance — typically between one and ten per cent. NeuroFrame has been validated against recognised tests and real KPIs, and the cognitive and personality domains are measured together in a single 30–60 minute sitting.

All Science section materials