Scientific foundations of the method

NeuroFrame has no personality theory of its own, and that is deliberate. The method is assembled from models that have been tested for decades, independently of us: four describe what is stable in human behaviour, two describe how to build an assessment so that observation counts as evidence.

What follows is what we take from each model, and in what capacity — plus a separate note on the professional testing standards our practice is aligned with.

Big Five: an interpretive frame, not a questionnaire

The Big Five is not one researcher's theory but the outcome of an empirical search that has been running since the 1930s. It began from a simple assumption: if a difference between people genuinely matters, language will have named it. Thousands of person-descriptive words were extracted from dictionaries, and for decades researchers examined how those descriptions clustered in real data. By the 1990s five factors — openness to experience, conscientiousness, extraversion, agreeableness, neuroticism — were replicating consistently across languages, cultures and sample types. It is that replication, rather than the elegance of the idea, that made the model the common language of differential psychology.

What NeuroFrame takes. We use the five-factor model as an interpretive frame, not as a questionnaire. Nobody answers a single question about themselves: the person solves a task, and the system records how they go about it. The Big Five enters at the next step — to establish which stable dimension of personality the recorded behaviour belongs to, and to name it in the terms the field already uses.

The distinction is not a matter of wording. A repackaged questionnaire would inherit the central weakness of self-report: people describe themselves as they see themselves, or as they would like to appear. Here behaviour comes first and the personality model arrives at interpretation. That also sets a limit on what we claim: our parameters relate to the five dimensions but are not identical to the scales of any particular questionnaire, and validity coefficients obtained for those questionnaires cannot simply be transferred to them.

RST: reinforcement sensitivity and the nature of risk

Reinforcement Sensitivity Theory describes not traits but brain systems that regulate behaviour towards reward and threat. Jeffrey Gray proposed it in the 1970s as a neuropsychological alternative to purely descriptive models of temperament. The current version is the revision by Gray and McNaughton (2000), which distinguishes the behavioural approach system (BAS), responsive to cues of possible gain; the fight-flight-freeze system (FFFS), which governs avoidance and defence; and the behavioural inhibition system (BIS), which engages when goals come into conflict and shifts behaviour into cautious checking.

What NeuroFrame takes. RST underpins how we interpret the risk appetite parameter. Risk here is not a synonym for courage or a character trait in the everyday sense, but an observable balance between two sensitivities: to potential gain and to potential loss. Behaviourally it shows in the conditions under which someone will accept a calculated risk for a future advantage, the conditions under which they prefer to preserve a secure position, and how that balance shifts as load and uncertainty increase.

One consequence matters when reading a report: a high value here is not "better" than a low one. A cautious balance is an advantage where the cost of error is high and a constraint where action is required before full information arrives. The parameter is compared with the requirements of a role, never with other people.

PEN: a biologically grounded view of temperament

Hans Eysenck developed his model from the late 1940s and settled on three dimensions: extraversion, neuroticism, and psychoticism, added in the 1970s — hence the abbreviation PEN. What set it apart from purely statistical approaches is that Eysenck looked for physiology behind the traits from the outset: differences in baseline arousal of the nervous system and in reactivity under load. Some of his specific neurophysiological explanations were not supported by later evidence, but the underlying move — looking for a stable biological mechanism behind stable behaviour — remains foundational to modern models of temperament, RST among them.

What NeuroFrame takes. PEN is coarser than the Big Five in resolution, and we use it not as an alternative but as a second vantage point, where what matters is the durability of a characteristic over time rather than fine differentiation. Two dimensions do the work in interpretation: emotional stability — how far decision quality holds up as pressure rises — and social activity — how attention is divided between the task and interaction with others.

The third dimension we do not interpret. It remains contested in the literature and has no bearing on the work behaviour we describe; the method draws no conclusions about health, private life, or attributes of a person outside the work context.

CB5T: why behaviour can tell you anything about personality

Cybernetic Big Five Theory (Colin DeYoung, 2015) answers a question the Big Five itself leaves open: what a trait actually is, and why the factors number what they number. The answer: traits are parameters of a system that governs behaviour through feedback. Such a system sets a goal, acts, compares the result against expectation, and adjusts either the action or the goal itself. Cybernetics here is not a computer metaphor but a language for describing control: stable differences between people are stable settings of the regulator.

What NeuroFrame takes. This is the theoretical bridge between behaviour and trait — what makes it legitimate at all to infer personality from actions rather than from self-description. If a trait is a control parameter, its expression should be sought at specific points in the cycle: how a goal is formulated, what happens on discovering a discrepancy between expectation and result, at what point strategy changes rather than merely effort, and what becomes of the goal when it proves unreachable.

The assessment environment is built so that this cycle repeats dozens of times and under rising load. Each repetition yields an independent observation, so conclusions rest on the consistency of a pattern rather than on a single striking episode.

Evidence-Centered Design: how behaviour becomes evidence

Evidence-Centered Design is not a theory of personality but an engineering methodology for building assessments. Robert Mislevy and colleagues set it out at the turn of the 2000s as an answer to a familiar failing in tests: a task is devised, it looks sensible, and yet the link between what the person does in it and what the test claims to measure is never made explicit.

ECD requires a chain of three links, kept open to scrutiny: the construct model — what exactly is being measured; the evidence model — which observable behaviour counts as evidence, and how strong that evidence is; and the task model — how to construct a situation in which that behaviour has any chance of appearing. If even one link goes unstated, the assessment rests on the designer's intuition.

How that looks here — types of behaviour and the parameters they serve as evidence for

  • Picking up unfamiliar rules quickly, moving from trial attempts to consistently sound decisions — learnability
  • Thinking a move through in advance and adjusting the plan on spotting a weak point before it causes a failure — progress monitoring
  • Exploring the available options independently and trying different approaches rather than one habitual tactic — openness
  • Holding decision quality steady as load increases, and changing approach after a setback rather than giving up — persistence

Stealth assessment: measurement embedded in activity

Valerie Shute described stealth assessment (Shute, 2011) — assessment embedded in an activity so that it is not experienced as a test. The idea is straightforward: if the environment already generates a continuous stream of actions, that stream can be accumulated as evidence and the estimate refined gradually, instead of running a separate procedure with a form and a stopwatch. In substance it is ECD carried through to practical form in a digital environment.

Why it matters. The test as a genre changes behaviour in itself: test anxiety sets in, along with guesses about the "right" answer and the wish to come across well. In questionnaires, candidates scoring systematically higher than incumbent employees is a robustly replicated effect, particularly on conscientiousness and emotional stability. When someone is absorbed in a task, much of that layer falls away.

What NeuroFrame takes. The format: 30 to 60 minutes in a game environment, unsupervised, with behaviour recorded as it happens. And a principle that goes with it: what is unobtrusive is the act of measurement, not the fact of assessment. The person knows they are being assessed, knows who commissioned it and which parameters will appear in the report. Covert observation of people who are unaware of it is neither part of the method nor permitted by it.

A fair caveat about the maturity of the approach: the empirical base for game-based assessment is still accumulating, and validity estimates depend on the design of the particular task. Each task is therefore validated in its own right rather than inheriting figures from the approach as a whole.

AERA/APA/NCME and ITC: alignment with requirements, not certification

The Standards for Educational and Psychological Testing are a joint document of three professional associations — AERA, APA and NCME — with the current edition published in 2014. They are not law but the de facto standard of the field: a body of requirements covering how validity is argued, how norms are built and documented, and how test security is handled. The International Test Commission (ITC) issues practical guidance in the same area, including its Guidelines on Security of Tests, Examinations, and Other Assessments.

Our practice is aligned with those requirements in two respects above all. Normative interpretation: a result is read only relative to a comparable group, never against an absolute scale — which is why a comparison base of more than 10,000 executives and entrepreneurs is what makes any percentile meaningful in the first place. Test security: the operational key is withheld from those being assessed, since otherwise an implicit measure turns into an explicit one and loses its value for everyone assessed thereafter.

The wording matters here. This is alignment with requirements, not certification: neither AERA/APA/NCME nor the ITC certifies or accredits assessment instruments — no such procedure exists — and we claim no certification of any kind. The one verifiable claim is this: the procedures are built to meet those requirements, and we are ready to show a client's own methodologists exactly how.

The models and their primary sources

ModelAuthorsYears
Big FiveAllport & Odbert · Tupes & Christal · Costa & McCrae1936–1992
RSTJ. A. Gray · Gray & McNaughton1970–2000
PENH. J. Eysenck · Eysenck & Eysenck1947–1976
CB5TC. G. DeYoung2015
Evidence-Centered DesignR. J. Mislevy · L. S. Steinberg · R. G. Almond2003
Stealth assessmentV. J. Shute2011
Standards · ITC GuidelinesAERA, APA, NCME · International Test Commission2014

These references are given so that claims can be checked; they imply neither affiliation with the authors nor any endorsement of the method by them.

Common questions

What is the method actually based on?

On six models, none of them ours. Four describe what is stable in human behaviour: the Big Five, Reinforcement Sensitivity Theory (RST), Eysenck's PEN model, and Cybernetic Big Five Theory (CB5T). Two describe how to build an assessment: Evidence-Centered Design supplies the chain of construct, evidence and task, while stealth assessment (Shute, 2011) supplies a format in which measurement is embedded in the activity. Our practice is additionally aligned with the Standards for Educational and Psychological Testing (AERA/APA/NCME) and with ITC guidance.

Is this just the Big Five repackaged?

No. A questionnaire asks people about themselves; NeuroFrame asks no such question — the person solves a task and the system records behaviour. The five-factor model enters at interpretation: it is what allows recorded behaviour to be assigned to a stable dimension of personality and named in accepted terms. The corollary is worth stating too: our parameters relate to the five dimensions but are not identical to the scales of any given questionnaire, so validity coefficients obtained for those questionnaires cannot be carried across.

What is Evidence-Centered Design?

An engineering methodology for building assessments, set out by Mislevy and colleagues at the turn of the 2000s. It requires an explicit chain of three links: what is being measured (the construct), which observable behaviour serves as evidence for it, and how to construct a situation that elicits that behaviour (the task). The point is auditability: if a link goes unstated, the assessment rests on the designer's intuition. The full evidence model for NeuroFrame is shared with a client's methodologists under a confidentiality agreement; the operational key — the exact feature weights and the model that converts behaviour into scores — is withheld from everyone, because disclosing it would devalue the normative base.

Is NeuroFrame certified by AERA/APA/NCME or the ITC?

No, and no instrument can be: these bodies publish standards and guidance but do not certify or accredit assessment methods — no such procedure exists. We speak of alignment with requirements rather than certification, and claim only what can be checked: normative interpretation rests on a comparison sample of more than 10,000 executives and entrepreneurs, and test security is maintained by withholding the operational key from those being assessed. We are ready to walk a client's methodologists through exactly how both work.

All Science section materials