Comparison
One Long Simulation or a Battery of Mini-Games
The difference between a long simulation and a battery of short mini-games is real and checkable — but it is not a difference in predictive accuracy, and we have found no study that compares the two formats head to head on one sample against one criterion. Where comparable evidence does exist, it points against the intuition that longer and more realistic means more accurate. What the long format actually changes is what enters the data: a battery resets state between tasks by construction, a continuous session carries it forward. This page sets out both architectures, the market as it stands on 4 August 2026, the findings that support the long format, the findings that refute claims we ourselves used to make, and the cases where a long session is the wrong instrument.
Updated:
Side by side
| Attribute | Mini-game battery | NeuroFrame |
|---|---|---|
| How the session is structured | Criteria describes its game-based assessments in its own words as "short, fun, and engaging mini-games" — several discrete tasks, each scored on its ownSource | One continuous session, no modules; state carries from the first minute to the last |
| Number of tasks in the battery | Test Partnership MindmetriQ: 6 named games — Net the Numbers, Number Racer, Link Swipe, Word Logic, Pipe Puzzle, Shape SpinnerSource | One task; no named sub-tasks |
| Length of a single task | MindmetriQ: 4–6 minutes per gameSource | Not applicable — the task is the session |
| Total published duration, battery | MindmetriQ in three named configurations: quick screening ~12 min, targeted selection ~18 min, comprehensive assessment ~35 minSource | 30–60 minutes, single configuration |
| Total published duration, another battery | Equalture: "a full assessment session usually takes between 15 and 45 minutes, depending on the number of games used for a role", and each game itself "lasts only a few minutes"; we found no library size on that pageSource | 30–60 minutes; the set does not vary by role — the role enters at the benchmark stage, not the task stage |
| A short published battery configuration in this snapshot | Criteria Cognify: three games, "just 10 minutes to complete, plus tutorials", measuring problem solving, numerical reasoning, verbal knowledge; the same page states verbatim that "a six-game version is also available on request"Source | No short configuration exists; at the top of a high-volume funnel this is a disadvantage on cost per completed assessment |
| Gamified modules inside a mixed battery | Sova Early Careers: 25 minutes, 7 components; the personality questionnaire and SJT are required, the two gamified challenges are marked optionalSource | No questionnaire component at all; all eight parameters are inferred from behaviour in the simulation |
| Nearest single-session products, published duration | Sova Immerse: one role simulation in three stages — manager briefing, task simulation, manager debrief — "15–30 min"Source | 30–60 minutes; a different genre — strategy simulation with resource constraint rather than role play |
| Nearest single-session products, published duration (2) | CapsimInbox: one inbox simulation, "typical inbox simulations are completed in just 15–60 minutes"; Hiring & Selection listed among use casesSource | 30–60 minutes; a different genre — no correspondence handling, decisions act on a simulated system |
| Single narrative; no duration found on the pages checked | Owiwi: a single narrative adventure, "Isles of the Shroud", rather than a set of mini-games; we found no stated duration on the product page, in the FAQ or in the technical manual as of 4 August 2026Source | 30–60 minutes, published; format is comparable here, length is not |
| Norm sample as published | Owiwi technical manual, page 21: "The norming sample is comprised of 5371 participants all over the world" — the short four-island version, completed from 2017 onward; the same page states that norming for the next four islands is in progress (vendor technical manual, PDF, p. 21; checked 4 August 2026)Source | 14,850 people from 500 companies, 21 industries, 21 functions, 8 grades. The two samples were built for different constructs, different populations and different purposes — these two cells are not a comparison of sizes. Our sample is smaller than that of instruments with forty-year histories, and that is on our list of weaknesses |
| Independent peer-reviewed study of the instrument | Owiwi: Nikolaou, Georgiou & Kotsasarlidou (2019), The Spanish Journal of Psychology — phase 1 N=193, phase 2 N=120; incremental validity over cognitive and personality tests demonstrated for academic performance; job performance measured by self-reportSource | None. We have no peer-reviewed publication of our own validation work — this is on our list of weaknesses, not a footnote |
| Reliability figure published by the vendor (different statistics — not directly comparable) | Test Partnership publishes, as its own figures, person reliability above 0.70 for MindmetriQ plus correlations with the ICAR battery of 0.91 verbal and 0.89 numerical (vendor's statement, checked 4 August 2026); we found no sample sizes or study design on that pageSource | Cronbach's alpha .69–.77 (below the conventional .80) and test–retest above .83. A person separation reliability and an alpha measure different things; do not read the two cells as a ranking |
Two classes of instrument, not one vendor. "Mini-game battery" and "one continuous simulation" are generalisations about format. Every product referred to here is named in full — Test Partnership MindmetriQ, Criteria Cognify, Equalture, Sova Early Careers, Sova Immerse, CapsimInbox, Owiwi — with a link to the vendor's own page and the date we read it, 4 August 2026. The section on search demand additionally quotes literal search queries that contain other companies' names; those are our own measurement of what people type into a search box, not statements about those companies' products. The names and trademarks belong to their respective owners and appear here nominatively, to identify the product itself, without logos or trade dress. We rate no product.
The short answer: the difference is real, and it is not about accuracy
There is no evidence that a long simulation predicts job performance better than a short battery. We have found no finding in the peer-reviewed literature that establishes it, and no head-to-head study of the two architectures — one continuous simulation against a set of independently scored mini-games, same sample, same criterion — in either direction. That absence is the first thing anyone being sold this distinction should know, and it is the first thing we say.
Where comparable evidence does exist, it runs against the intuition that immersion buys accuracy. In the meta-analytic recalibration by Sackett, Zhang, Berry and Lievens (2022), high-fidelity methods sit below methods with no immersion at all: work sample tests at .33 and assessment centres at .29, against structured interviews at .42 and job knowledge tests at .40. Lievens and Patterson (2011) found that both a high-fidelity assessment centre and a low-fidelity situational judgement test added incremental validity over knowledge tests — fidelity itself is not the driver.
So what does change. The architecture of observation changes. A battery resets state between tasks: each task starts clean, ends, and is scored on its own. A continuous session carries state forward — what a person did in minute ten becomes the condition of minute forty. In one architecture the observable "how does this person handle the consequences of their own earlier decision" exists; in the other, by our definition of a battery, it does not — not because it is measured badly but because there is nothing for it to attach to. That is a claim about what enters the data, not about how accurately the data predicts.
Below: definitions of both formats, a market snapshot dated 4 August 2026 with links to vendor pages, what the literature supports, the four claims we deleted from our own materials because the literature refutes them, and a section on where a long simulation is the wrong instrument.
Two terms that do not exist in the field: mini-game battery and single-simulation assessment
The existing taxonomy in the literature classifies assessments by how game-like they are. Landers and colleagues (2022) distinguish game-based, gamified and gamefully designed assessments and set out what validating each requires. That axis is useful, and it says nothing at all about session architecture: a gamefully designed instrument can be six separate tasks or one, and the label does not tell you which.
So we introduce a second axis, orthogonal to the first. These two names are ours, coined on this page. They are not established terminology, and we will not cite them as though they were.
A mini-game battery is an assessment assembled from several short tasks, each with its own rules, its own beginning, its own end and its own score. State does not carry between tasks. The set is configurable per role, and the total duration is the sum of its parts — remove a module and the assessment is shorter but otherwise intact.
A single-simulation assessment is an assessment in which the whole session is one task, with one rule set and one persisting state. The outcome of early actions becomes an input to later ones. The session cannot be decomposed into independently scored parts without destroying that state, and its duration is not a sum of modules because there are no modules.
The boundary is not duration and not degree of gameness. A thirty-five-minute battery is still a battery; a fifteen-minute continuous role simulation is still one simulation. Test Partnership publishes MindmetriQ in a configuration of roughly thirty-five minutes made of six named games of four to six minutes each — long in total, discrete by construction. Sova publishes Immerse as a single continuous role simulation of fifteen to thirty minutes in three stages — shorter in total, continuous by construction. Duration and architecture are independent properties, and conflating them is the most common error in this discussion.
One more distinction worth keeping separate: continuity is not the same as narrative framing. A set of independent tasks can be wrapped in a single story and still reset state between them. What the definition above turns on is whether earlier behaviour changes later conditions, not whether the screens share a fictional world.
What the market offers: a snapshot dated 4 August 2026
Everything in this section is taken from vendor pages, read on 4 August 2026, and stated as the vendor states it. We give no assessment of any product.
Criteria describes its game-based assessments in its own words as "short, fun, and engaging mini-games". Its product Cognify is published as three games taking about ten minutes plus tutorials, measuring problem solving, numerical reasoning and verbal knowledge. Criteria also states on that page that Cognify was independently validated by two studies, one by a multinational technology company and one by the psychologist Dr Richard Landers. We found no numerical coefficients and no sample sizes on that product page, and we did not look for them in the vendor's other materials.
Test Partnership publishes MindmetriQ as six named games — Net the Numbers, Number Racer, Link Swipe, Word Logic, Pipe Puzzle, Shape Spinner — of four to six minutes each, offered in three configurations: quick screening at about twelve minutes, targeted selection at about eighteen, comprehensive assessment at about thirty-five. On the same page the vendor publishes its own correlations with the ICAR reference battery: 0.91 for verbal and 0.89 for numerical reasoning, a range of 0.75 to 0.84 across the battery, and person reliability above 0.70 (the vendor's figures, checked 4 August 2026). We found no sample sizes and no description of the study design on that page; we did not look for them in the vendor's other materials.
Equalture states in its FAQ that a full session usually takes between fifteen and forty-five minutes depending on how many games a role calls for, and that each game itself lasts only a few minutes. We found no total for the number of games in the library on that page; we did not look for one in the vendor's other materials. Sova publishes its Early Careers assessment as twenty-five minutes across seven components, of which the personality questionnaire and the situational judgement test are required and the two gamified modules — Gamified Numerical Challenge and Gamified Pattern Challenge — are marked optional.
Products built as one continuous session exist and are published as such. Sova Immerse is described as a single role simulation in three stages — manager briefing, task simulation, manager debrief — lasting fifteen to thirty minutes. CapsimInbox is described as one inbox simulation completed in fifteen to sixty minutes, with Hiring & Selection listed explicitly among its use cases. Owiwi is built as a single narrative, "Isles of the Shroud", rather than a set of separate mini-games; we found no stated duration on its product page, in its FAQ or in its technical manual as of 4 August 2026, so the format is comparable here and the length is not.
A necessary caveat about the limits of this check. We read a selection of English-language vendor pages on one date. That is not a survey of the market, so we do not write "no one", "only" or "every vendor". What can be said is narrower and still useful: among the products we checked, the prevailing published architecture is a battery of short tasks, and continuous single-session products are published at durations of fifteen to sixty minutes rather than as a thirty-minute-plus minimum.
What the science says about the game format itself
Before the two architectures can be compared, the honest state of the evidence on game-based assessment as a whole has to be stated, because it sets the ceiling for anything either format can claim.
The most direct source is the PRISMA systematic review by Ramos-Villagrasa, Fernández-del-Río and Castro (2022), covering thirty-four articles. Its conclusion is uncomfortable for anyone selling game-based assessment, including us: game-related assessments do not offer sufficient advantages to recommend their use over conventional methods, unless improving applicant reactions — organisational attractiveness in particular — is thought to offer added value. Construct validity is described as inconclusive. Evidence on faking is described as too scarce to draw conclusions. Adverse impact is not worse than conventional tests. Candidate reactions are the one area with a confidently positive result, and even there the review documents a game-framing phenomenon: simply calling a test a game improves reactions without changing its content.
Two meta-analyses map game-based measures onto traditional ones. Bipp, Wee, Walczok and Hansal (2024) pooled fifty-two samples from forty-four studies and more than six thousand one hundred adults: the relationship between game-related scores and traditional cognitive ability tests was r = .30 observed and .45 corrected, and games specifically designed to measure cognitive ability did not correlate more strongly than the rest. The authors leave open whether game-related assessments measure the same thing traditional tests do. Fadillah, Hidayat and Santoso (2025) pooled eighteen studies from thirteen articles on convergent validity against self-report personality measures, r = .516, and name circular validation and the absence of standardised frameworks as limitations of that body of work.
The gap that matters most is what none of these cover. We have found no meta-analysis of the criterion validity of game-based assessment against job performance. Every meta-analysis we did find is about convergence — agreement with traditional instruments — not about predicting an outcome. A predictive claim for this class of instrument — ours or anyone else's — rests on the data of whoever makes it, because we found no literature for it to rest on.
This is the ceiling. Within it, the defensible position for a game format is: comparable validity to conventional methods, materially better candidate reactions, adverse impact no worse. Everything argued below about session architecture sits inside that ceiling and does not raise it.
Four claims we deleted from our own materials
This section exists because a comparison published by an interested party is worth reading only if it names the places where the evidence went against the party publishing it. Each of the four claims below appeared in our materials. Each has been removed.
One: "at around minute five to seven the brain stops treating the session as a test and enters flow." We found no such timing in the literature we reviewed. Durcan, Holland and Bhattacharya (2024) state directly that little attention has been paid to the time required for flow states to emerge once an activity has started. The estimates that do exist differ by two orders of magnitude: Kotler, Mannino, Kelso and Huskey (2022) model flow onset within the first seconds; a study by Yun and colleagues, cited in that review, had expert gamers reporting they needed at least twenty-five minutes; Genc and colleagues (2026) find flow rising in the third to fifth epoch of a twenty-five-minute session. The figure of five to seven minutes appears in none of them. We no longer name a minute at all.
Two: "immersion switches off social masks." Loss of self-consciousness is one of the constituting features of flow in Csikszentmihalyi's definition (Flow: The Psychology of Optimal Experience, Harper & Row, 1990) — part of what the construct means, not an independently demonstrated consequence of it. We found no study in the game-based assessment literature that measures flow as a mediator of reduced socially desirable responding. The mechanism is discussed as a hypothesis, including in Ohlms and colleagues (2025), and it has not been tested directly. What can be said instead is narrower and still true: a game records what a person does rather than what they say about themselves, so there is no obvious right answer to select.
Three: "game-based assessment is more accurate and more honest than questionnaires." The systematic review of thirty-four studies concludes the opposite, as quoted above. The defensible version is comparable validity with better candidate reactions — and that is what we now write.
Four: "a long simulation yields higher predictive validity." The numbers point the other way. In Sackett and colleagues (2022), high-fidelity work samples sit at .33 and assessment centres at .29, both below structured interviews at .42 and job knowledge tests at .40. Lievens and Patterson (2011) showed that both low-fidelity and high-fidelity simulations add incremental validity over knowledge tests, which is precisely the finding that removes fidelity as an explanation. Our argument for a continuous session is now about the nature of the data, not about superior prediction.
A fifth item belongs here for completeness, and it concerns a metric rather than a sentence: we do not report an individual "fatigue index". The reasoning is in the section on time on task.
What a long session genuinely brings: a partly different construct
The scientifically defensible case for a dynamic session with delayed consequences is not that it measures better. It is that it measures something partly different.
The literature on complex problem solving in dynamic microworlds is the right frame. Stadler, Becker, Gödker, Leutner and Greiff (2015) meta-analysed forty-seven studies and found a weighted mean correlation between complex problem solving and intelligence of r = .43, with a 95% confidence interval of .37 to .49. That corresponds to roughly eighteen per cent shared variance. Complex problem solving overlaps with general cognitive ability substantially and does not reduce to it — which is exactly the shape of claim a continuous simulation can make.
The honest qualifications go with it. There is evidence of incremental validity for complex problem solving over intelligence for academic and work outcomes, and the field is contested: a portion of the literature holds that the evidence is not yet sufficient to recognise complex problem solving as a distinct ability. The meta-analysis is about microworld tasks in research settings, not about commercial hiring instruments. And we have not published evidence that our own simulation measures complex problem solving specifically; we cite this literature as the frame that makes the architectural argument coherent, not as a validation of our product.
Stated carefully, the claim is this. A dynamic environment in which decisions have delayed consequences engages a construct that classical tests capture only in part. That is a reason to consider such an environment as a complement to conventional measurement — and it is not a reason to expect a higher validity coefficient, because the meta-analytic numbers on high-fidelity methods say otherwise.
Observation under load — and why we do not publish a fatigue index
The vigilance decrement — the decline in sustained attention and in the quality of responses as time on task increases — is one of the most reproducible effects in cognitive psychology. Typical vigilance tasks in research run thirty to forty-five minutes, and the decline is measurable within the first ten. Mental fatigue also shifts decision parameters, including a documented move toward risk aversion. This is a group-level effect and it is well established.
What follows from it for assessment design is modest and worth stating precisely. A session long enough for load to accumulate lets an observer see behaviour in both a fresh and a loaded state. A set of short tasks, each starting clean, largely records the first. That is a difference in what is available to observe.
What does not follow is a personal metric. A score of the form "how much did this individual's performance drop by the end of the session" is a difference score, and difference scores are notoriously unreliable: the difference between two quantities each measured with error is almost always less reliable than either. Hedge, Powell and Sumner (2018) documented the reliability paradox directly — tasks producing robust group effects showed test–retest reliabilities ranging from zero to .82 and were in most cases unsuitable for measuring individual differences, precisely because between-person variance in them is small.
So we say what the format allows us to observe, and we do not claim a validated fatigue scale. If any such indicator ever appears in a report an employer receives, its test–retest reliability has to appear next to it. Anyone comparing instruments should apply the same rule to every vendor, including this one: ask what the reliability of the specific number in the report is, not the reliability of the instrument in general.
Engagement across a session, and deliberate distortion
Two findings sit under the long format that are worth having, provided they are stated with their limits.
Engagement rises across a session rather than being set in the first minutes. Genc, Surer, Wittmann, Popov and Lenggenhager (2026) measured flow continuously — participants held a pedal — across a twenty-five-minute VR session divided into five five-minute epochs. Press frequency was higher in the third, fourth and fifth epochs, and post-session questionnaires correlated more strongly with the final five minutes than with the earlier ones. The practical reading is that a longer format yields more data collected while engagement is high. The limits are real: a small sample, a VR game rather than a hiring assessment, and the authors themselves name the peak-end rule as an alternative explanation for the questionnaire correlations. This finding supports "engagement rises across the session"; it does not support any statement about a specific minute.
On deliberate distortion, Ohlms, Melchers, Kanning and Barends (2025) ran a controlled experiment with 171 participants randomly assigned to answer honestly or as an applicant would, taking both a game-based measure and a conventional honesty-humility test. Participants were able to distort their results in both, and the faking effect in the game-based measure was significantly smaller. The authors' own conclusion is that game-based assessments are not a panacea to prevent it completely. A counterweight belongs here: the 2022 systematic review located only one faking study in the whole field and called the evidence too scarce to draw conclusions, which leaves the Ohlms result a single experiment rather than a body of evidence.
So the supportable sentence is: a game format reduces, and does not remove, socially desirable responding — and this rests on a single controlled experiment, not on a body of evidence. Note also that this is a property of the game format as such, not of session length. We found nothing in this literature that distinguishes a long simulation from a short battery on faking, and we do not claim that it does.
Delayed consequences: a difference by construction, not by degree
Here is where the two architectures differ most sharply, and it is worth being precise about what kind of difference it is.
By our definition, a battery does not carry state between tasks: each task starts from a defined initial state and is scored on its own. The definition is ours. Whether state carries inside any particular product is something only its vendor knows, and we found no description of that mechanism on the pages we read on 4 August 2026 — so this is a statement about an architecture, not an attribution to the products named above. In a continuous session, the state at minute forty is a product of what the candidate did at minute ten. A resource spent early is not available later. Effort saved early turns into a problem that has to be handled later.
The consequence is a class of observable. "How does this person deal with a situation they themselves created, without being told that they created it" is a thing that can be recorded in one architecture and cannot be recorded in the other. Not measured less accurately — not present. That is a structural statement, and it can be checked against any vendor's own description of its product rather than needing a study.
And now the limit on that statement, which is as important as the statement. This does not establish that the resulting data predicts work outcomes better. There is no head-to-head research comparing a single long simulation with a battery of mini-games on the same sample against the same criterion — none that we found, in either direction. If such a study were run and came out against the long format, that would be a real result and we would publish it here. Until then, the correct form of the argument is: the two architectures make different behaviour available to observation, and which of them predicts better is an open empirical question.
One further limit, on us specifically. Continuity buys observation and costs coverage. A single continuous task can only measure what that task calls for. It does not have modules to add when a role needs verbal or numerical ability measured directly, and it does not have an alternate form for retesting. A battery has both by construction. Those are not small trade-offs, and the section on where a long simulation is the wrong instrument takes them seriously.
The situation explains as much as the competency — and that cuts both ways
A well-replicated finding in the assessment centre literature is that exercise factors account for at least as much variance in ratings as the dimensions the exercises are supposed to measure. Bowler and Woehr (2006) established this meta-analytically, following the earlier work of Lievens and Conway (2001), Journal of Applied Psychology 86(6), 1202–1222, doi:10.1037/0021-9010.86.6.1202. Behaviour, in other words, is strongly situational: what a person does depends heavily on the specific situation they are placed in, not only on a stable trait they carry between situations.
One reading of this supports a continuous scenario. If behaviour is situational, then observing a function stripped out of context changes what is being observed. A whole scenario at least presents behaviour bound to a context rather than as an isolated response.
The other reading cuts against us, and we would rather say it ourselves than have a methodologist say it for us. If exercise variance dominates dimension variance, that is an argument against any claim to be cleanly measuring a dimension inside a simulation — including ours. A continuous scenario is a very large single exercise. It does not escape the exercise effect; if anything, it is the exercise effect. The eight parameters we report are inferences from behaviour inside one specific situation, and the extent to which they generalise beyond that situation is an empirical question about our own instrument, not something the assessment centre literature settles in our favour.
This is why we describe our own numbers as percentiles against a comparison sample, published together with their internal consistency of .69 to .77 and test–retest reliability above .83, rather than as trait measurements. And it is why the standard caution attached to every range we publish reads: a gap on one or two parameters is something to explore at interview, not grounds for rejection.
The preparation industry: a measurable property of discrete named tasks
A named, discrete task can be found, described and rehearsed. That is not an opinion about any product; it is a property of the format, and it leaves a measurable trace in search demand.
Our own search-demand measurement, taken on 4 August 2026 in Semrush: in the UK database, "arctic shores practice test free" runs at 390 searches a month, and five queries of the form "arctic shores <name> game" — Balance, Order, Predict, Lock and Ticket — come to 550 a month between them (170, 110, 110, 90 and 70). In the US database, "pymetrics games answers" runs at 110. Those are the names the prep industry types into a search box; we do not attribute them to any vendor, and we make no claim here about how any vendor names or describes the parts of its own product. We give no link for these figures because they are not a vendor statement but a demand measurement; the measurement reproduces in any keyword tool, and we name the tool, the database and the date so it can be re-run and checked.
What follows from this, and what does not. It follows that when a task has a public name, an ecosystem of guides forms around that name — a structural consequence of naming and discreteness. It does not follow that preparation works, that scores shift because of it, or that any instrument is worse for it. We have no data on the effect of preparation on final scores, and we make no such claim. Vendors also respond to this: item pools, randomisation and implicit scoring are standard countermeasures, and a search volume figure says nothing about how effective those countermeasures are.
The symmetrical point about ourselves. Our session is one task with no publicly named modules, so today there is nothing granular to search for. That is a property of our current design and not a virtue we earned — and it would disappear the moment we ever named modules. Exposure risk applies to us too, and more sharply than to a battery: we run a single task with no alternate form for retesting, which is on our own list of weaknesses. We also do not publish the scoring key — the feature weights, the behaviour-to-score model, or any reference path through the simulation. That is a test security requirement under the AERA/APA/NCME standards and the ITC guidelines on test security, and it applies to every serious instrument rather than being a distinguishing feature of ours.
What this means when you are choosing an instrument
Start from the task, not from the format. The architecture question is downstream of what you are actually trying to observe.
High-volume screening at the top of the funnel on cognitive ability. A short battery fits this: ten to fifteen minutes, one module per construct, configurable per role, and a completion cost that scales. A thirty-to-sixty-minute session at this stage buys observation you do not need and pays for it in drop-off.
Verbal or numerical ability specifically. A battery has dedicated modules for these. A single continuous simulation does not, and ours does not — we do not cover verbal and numerical ability, and we say so before a pilot rather than after.
Behaviour under the accumulating consequences of a person's own decisions. This is the case a continuous session is for. It is also the case where you should ask hardest for evidence, because we found no meta-analysis of criterion validity to fall back on and everyone, ourselves included, will be citing their own data.
The cheapest available improvement in prediction. If the budget allows one change, structuring the interview is probably it: .42 in the Sackett recalibration, against .33 for work samples and .29 for assessment centres. Buying any game format before the interview is structured is optimising the wrong stage. We would rather say this and lose a deal than have a methodologist work it out after signing.
Candidate experience. This is where the game format has its most consistent positive evidence — and read it with the game-framing caveat from the systematic review, since reactions improve from calling something a game even when the content is unchanged.
Whatever you choose, ask every vendor the same five questions and compare the answers rather than the marketing. What is the sample size and composition of the norm group? What are the internal consistency and test–retest reliability of the specific scores in the report? Is there an adverse impact report, with sample and date? Who performed the validation, and is anything peer-reviewed? And for any indicator built as a change over time, what is its reliability? Applied to us, those questions have uncomfortable answers, which are in the next section.
Where a long simulation is the wrong instrument
This is about the format, not about any one product, and it is the section a page published by an interested party usually leaves out.
Throughput. Thirty to sixty minutes against ten costs completion. Criteria publishes Cognify at about ten minutes plus tutorials; Sova publishes an Early Careers battery at twenty-five. At the top of a funnel of thousands, the shorter instrument wins on cost per completed assessment, and no argument about observational richness changes that arithmetic.
Construct coverage. A battery adds modules; a continuous session does not. If a role genuinely turns on numerical reasoning, a battery measures it directly and a simulation infers it at best indirectly.
Retesting. Batteries can substitute forms between sittings. A single continuous task has no alternate form, which limits repeat use inside one organisation.
Configurability. A battery lets a client drop a module that a role does not need. A continuous session is all or nothing: you cannot remove the middle twenty minutes.
Accessibility. A long session in real time under rising time pressure is a harder format for candidates with motor or visual limitations than an untimed questionnaire, and it carries a correspondingly sharper risk of indirect discrimination. Any vendor selling this format should be asked for its adaptation policy in writing. Ours is not yet formalised — we say so plainly, and until it is, we recommend not using the result as the sole basis for a decision and providing an alternative procedure for anyone the format does not suit.
Predictive validity as such. If the only criterion is the size of the coefficient, the recalibrated meta-analytic figures do not favour immersion: structured interviews at .42 and job knowledge tests at .40 sit above work samples at .33 and assessment centres at .29. A long simulation is a case to make about what you observe, not about the number you get.
Disclosure: who wrote this and how it was checked
This comparison is published by NeuroFrame — that is, by an interested party. Which is why every fact about someone else's product carries a link to its primary source, and why we assign no ratings at all.
How the vendor facts were gathered: we read vendor pages in English on 4 August 2026 and quote what they state, dated. We do not describe a market, because we did not survey one, and we avoid the words "only", "no one" and "every vendor" for that reason. Where we could not find a figure on the pages we read — a duration for Owiwi, a count of games in the library for Equalture — we say exactly that: we did not find it on those pages as of 4 August 2026. That is a statement about our check, not a statement that the figure does not exist or that the vendor does not publish it.
How the science was checked: the validity coefficients come from the Sackett, Zhang, Berry and Lievens (2022) recalibration read in the source article, not from secondary summaries, several of which circulate with wrong numbers. We do not use the 1998 Schmidt and Hunter figures as current values; the recalibration showed systematic overcorrection for range restriction across all five common approaches, lowering the estimates by .10 to .20.
What we changed as a result of writing this page: a flow-onset timing claim, a claim about immersion suppressing social masks, a claim that game-based assessment beats questionnaires, and a claim that a long simulation raises predictive validity — all four removed, with the sources that refute them named in the section above.
Regulatory context, described rather than interpreted. The main body of the EU AI Act became applicable on 2 August 2026, and systems used in recruitment and selection fall under Annex III as high-risk. Under New York City Local Law 144 the duty to publish a bias audit lies with the employer, not the vendor; a vendor audit is a voluntary signal, and an auditor connected to the development, employment relationship or financial interest does not count as independent. We are not giving legal advice here — check the primary sources and your own counsel.
Our own numbers on this page come from a single internal source of record: session duration 30–60 minutes, comparison sample 14,850 people across 500 companies, 21 industries, 21 functions and 8 grades, internal consistency .69–.77, test–retest above .83, confirmatory fit index .96, and a profile library of 287 professions across 34 occupational areas and 8 lifecycle stages, 2,296 benchmark profiles in total, each carrying its own human-readable rationale.
When to choose someone else
Not one peer-reviewed publication of our own validation studies. We cite other people's meta-analyses with DOIs; our own coefficients have not been through review or independent replication. If your procurement requires published validation, we do not meet it today.
No formal bias audit and no adverse impact report. We do not substitute the phrase "our method is objective" for the document that does not exist. What we have instead is architectural: we receive no name, no gender and no age, only codes the client distributes internally. That removes the risk of a protected attribute being used directly; it does not remove the indirect risk, because behavioural measures can correlate with gender, age or a health limitation, and the only way to check that is to measure adverse impact on the output. We have no such measurement yet, and a property of the design is not a substitute for one.
Internal consistency of .69 to .77, below the conventional .80 threshold expected of instruments used in high-stakes decisions. We compensate with test–retest reliability above .83 — stability over time rather than consistency of wording — and we do not hide the first number behind the second.
A comparison sample of 14,850 people, smaller than instruments with forty-year histories. Spread across 21 industries, 21 functions and 8 grades, individual cells hold hundreds rather than thousands. For narrow roles the range is a reference point, and a proper comparison needs calibration on your own people.
83% of the library's 2,296 records sit at specialist / individual-contributor level, and top management is covered by nine professions. If your problem is executive assessment, that is where our library is thinnest.
One task instead of a battery: we do not cover verbal or numerical ability, we have no situational judgement module, and we have no alternate form for retesting. We complement existing instruments; we do not replace them.
Our compliance perimeter is Russian. Federal Law 152-FZ is closed architecturally because we receive no personal data at all, but we have no GDPR package, no EU AI Act documentation set and no NYC LL144 materials. For an EU or US procurement that is a stop, and we say it before a pilot rather than after.
No ATS connectors — a direct consequence of the anonymised-code scheme, since ATS integration normally means exchanging exactly the data we choose not to hold. At high volume the manual step costs someone's time.
No formalised adaptation policy for candidates with motor or visual limitations. Until there is one, do not use the result as the sole basis for a decision, and provide an alternative route for anyone the format does not suit.
The client-facing handbook exists only in Russian. A multilingual workforce cannot be given a range rationale in a language most of them do not read.
Questions
- Which is harder to fake — one long simulation or a battery of mini-games?
- On the available evidence, neither. The game format as a class reduces deliberate distortion relative to a questionnaire: in a controlled experiment with 171 participants (Ohlms et al., 2025) the faking effect in the game-based measure was significantly smaller, though distortion was still possible, and the authors state plainly that game-based assessments are not a panacea. We found no study comparing long and short game formats on this. The 2022 systematic review found only one faking study in the entire field and called the evidence too scarce for conclusions.
- Does a long simulation predict job performance better?
- No, and there is no evidence that it does. We found no head-to-head comparison of the two formats on one sample against one criterion. The indirect evidence points the other way: in the Sackett et al. (2022) recalibration, high-fidelity methods — work samples at .33 and assessment centres at .29 — sit below structured interviews at .42 and job knowledge tests at .40, and Lievens and Patterson (2011) showed that low-fidelity simulations add incremental validity too. Fidelity itself is not the driver.
- At what minute does a candidate relax and start behaving naturally?
- We found no such minute in the literature, and we no longer name one. The review by Durcan, Holland and Bhattacharya (2024) states directly that little attention has been paid to how long flow takes to emerge, and the available estimates differ by two orders of magnitude — from the first seconds (Kotler et al., 2022) to twenty-five minutes or more by the self-report of expert gamers. What can be said: in continuous measurement, engagement is higher in the second half of a session than in its opening minutes (Genc et al., 2026 — small sample, VR game).
- Won't this just measure who plays more video games?
- That is a legitimate objection to any game-based assessment, and the field has no universal answer to it. In our design the rules are explained before the session starts: there is nothing to guess, what is measured is how well known rules are applied as load rises, and the speed of getting comfortable with an unfamiliar format is treated as a signal of learning agility rather than as an advantage. That is an argument about design, not published proof that prior gaming experience has no effect — we have no such proof.
- We need to screen five thousand candidates. What should we use?
- Most likely a short battery. At the top of the funnel, cost per completed assessment decides: Criteria publishes Cognify at about ten minutes plus tutorials, Sova publishes an Early Careers battery at twenty-five. A thirty-to-sixty-minute session at that point buys observation the task does not require and pays for it in drop-off. A continuous simulation belongs further down the funnel, at lower volumes.
- Is there a study that compares the two formats directly?
- We did not find one — neither for nor against the long simulation. We also found no meta-analysis of the criterion validity of game-based assessment against job performance; every meta-analysis we did find is about convergence with traditional methods. That is the honest state of the field as of 4 August 2026, and a vendor claiming a numerical advantage here would be citing its own data, since we found no literature to cite.
- We have budget for one process improvement. Where should it go?
- Most likely into structuring your interviews. In the Sackett et al. (2022) recalibration, structured interviews came out at .42, above work samples at .33 and assessment centres at .29. Buying any game-based instrument before the interview is structured optimises the wrong stage. It is against our interest to say this, but it is easy to verify and expensive to discover after signing.
- How does game-based assessment sit with the EU AI Act and NYC LL144?
- The main body of the EU AI Act became applicable on 2 August 2026, and recruitment and selection systems fall under Annex III as high-risk — that applies to any vendor with EU clients regardless of format. Under New York City Local Law 144, the duty to publish a bias audit lies with the employer, not the vendor; a vendor audit is a voluntary signal, and an auditor connected to the development, the employment relationship or a financial interest does not count as independent. We are not giving a legal opinion here — check the primary sources.
Sources
Every link was opened and checked on the date shown above.
- Revisiting meta-analytic estimates of validity in personnel selection — Sackett, Zhang, Berry & Lievens (2022), Journal of Applied Psychology 107(11), 2040–2068
- Revisiting the design of selection systems in light of new findings regarding the validity of widely used predictors — Sackett, Zhang, Berry & Lievens (2023), Industrial and Organizational Psychology 16(3), 283–300
- The validity and utility of selection methods in personnel psychology (historical source of the 1998 estimates) — Schmidt & Hunter (1998), Psychological Bulletin 124(2), 262–274
- Game-related assessments for personnel selection: A systematic review — Ramos-Villagrasa, Fernández-del-Río & Castro (2022), Frontiers in Psychology 13:952002
- The relationship between game-related assessment and traditional measures of cognitive ability: a meta-analysis — Bipp, Wee, Walczok & Hansal (2024), Journal of Intelligence 12(12):129
- Convergent validity of game-based assessment: a meta-analysis — Fadillah, Hidayat & Santoso (2025), International Journal of Serious Games
- Game-based, gamified, and gamefully designed assessments for employee selection: definitions, distinctions, design, and validation — Landers et al. (2022), International Journal of Selection and Assessment 30(1)
- A framework for neurophysiological experiments on flow states — Durcan, Holland & Bhattacharya (2024), Communications Psychology 2:66
- First few seconds for flow: a comprehensive proposal of the neurobiology and neurodynamics of state onset — Kotler, Mannino, Kelso & Huskey (2022), Neuroscience & Biobehavioral Reviews 143:104956
- Tracking flow in real time during a virtual reality gaming session — Genc, Surer, Wittmann, Popov & Lenggenhager (2026), Psychophysiology 63(4):e70283
- Game on, faking off? An experimental study of faking in game-based and conventional assessment — Ohlms, Melchers, Kanning & Barends (2025), Journal of Business and Psychology
- The reliability paradox: why robust cognitive tasks do not produce reliable individual differences — Hedge, Powell & Sumner (2018), Behavior Research Methods 50, 1166–1186
- Examining the role of task requirements in the magnitude of the vigilance decrement — PubMed Central — review of time-on-task effects and typical vigilance task durations
- Effects of mental fatigue on risk preference and feedback processing in risk decision-making — PubMed Central
- Dimension and exercise variance in assessment center scores: a large-scale evaluation of multitrait-multimethod studies — Lievens & Conway (2001), Journal of Applied Psychology 86(6), 1202–1222
- Standards for Educational and Psychological Testing (test security requirements) — AERA, APA & NCME (2014)
- A meta-analytic evaluation of the impact of dimension and exercise factors on assessment center ratings — Bowler & Woehr (2006), Journal of Applied Psychology 91(5), 1114–1124
- The validity and incremental validity of knowledge tests, low-fidelity simulations, and high-fidelity simulations — Lievens & Patterson (2011), Journal of Applied Psychology 96(5), 927–940
- Complex problem solving and intelligence: a meta-analysis — Stadler, Becker, Gödker, Leutner & Greiff (2015), Intelligence 53, 92–101
- Exploring the relationship of a gamified assessment with performance — Nikolaou, Georgiou & Kotsasarlidou (2019), The Spanish Journal of Psychology 22
- Criteria — Assessments catalogue ("short, fun, and engaging mini-games") — Criteria Corp, vendor page
- Criteria — Cognify (three games, about 10 minutes plus tutorials) — Criteria Corp, vendor page
- Test Partnership — Gamified Assessment (MindmetriQ): six games, 4–6 minutes each, three configurations — Test Partnership Limited, vendor page
- Equalture — FAQ (session 15–45 minutes, each game a few minutes) — Equalture, vendor page
- Sova Assessment — Early Careers (25 minutes, seven components, gamified modules optional) — Sova Assessment Limited, vendor page
- Sova Immerse — one continuous role simulation in three stages, 15–30 minutes — Sova Assessment Limited, vendor page
- Capsim — Inbox Simulations (15–60 minutes; Hiring & Selection listed as a use case) — Capsim Management Simulations, vendor page
- Owiwi — Product (single narrative "Isles of the Shroud"; no duration found on the page) — Owiwi, vendor page
- Owiwi Soft Skill Manual (technical manual, PDF): norm sample 5,371 on p. 21; construct validity N=938 on p. 17 — Owiwi, vendor technical manual
- The ITC Guidelines on the Security of Tests, Examinations, and Other Assessments (Final Version v1.0, 6 July 2014, ITC-G-TS-20140706) — International Test Commission (2014)
- EU AI Act — implementation timeline — artificialintelligenceact.eu