Comparison
Game-Based Assessment: Four Classes of Tools, and How the Vendors Differ
The word "game" on vendor sites covers at least four different things: a battery of three-minute mini-games, one continuous simulation, a package where a questionnaire carries most of the weight and two modules are gamified, and a plain self-report survey wrapped in a game interface. Which class you are buying determines what the tool can observe, how much of a candidate's time it takes, and how easy it is to rehearse. This page maps the four classes, puts nine vendors and ourselves into one table with a link to each vendor's own page, and states plainly where we lose. Every vendor figure here was checked on 4 August 2026 and is quoted from the vendor's own published material.
Updated:
Side by side
| Vendor | Class, format and duration, as the vendor publishes them | NeuroFrame |
|---|---|---|
| pymetrics (Harver) | Class: mini-game battery. Vendor page states "12+ interactive, gamified experiences" and "25 Minutes to complete". The peer-reviewed ACM FAccT 2021 paper, co-authored by Northeastern researchers and pymetrics staff, describes a "core set of twelve games" built on classical paradigms (BART, Trust Game, Go/No-Go, digit span, Iowa Gambling Task, Flanker, Tower) and states that models are trained per client, typically on 50–100 high performers. Checked 4 August 2026.Source | One continuous simulation, 30–60 minutes, no mini-game battery. The role benchmark is built externally — 287 professions from the Russian occupational register × 8 lifecycle stages = 2,296 records — not trained on the client's current high performers. |
| Arctic Shores | Class: battery of short interactive tasks. On the vendor's pages two wordings sit side by side: the marketing line "tasks, not questions" and the glossary entry "our game-based assessment". Vendor brochure: "a series of up to eight engaging tasks"; candidate guidance: "up to nine tasks"; How it works page: "know true potential, in 45 minutes". An employer's public candidate briefing we checked asks candidates to set aside 15–20 minutes, because the employer configures the task set. Checked 4 August 2026.Source | The session is not configured per role: everyone completes the same 30–60 minute simulation. What is configured per role is the comparison benchmark, and it is published as data — 2,296 records, each with a human-readable rationale (1.07 million characters in total). |
| HireVue | Class: mixed. The game block is a module inside the assessment platform; the vendor's ebook names five pre-built cognitive packages, three of them Cognition Short, Standard and Comprehensive, and states the games "take under 15 minutes to complete"; a per-package duration we did not find. Candidate page: "each game takes approximately 3 minutes", sessions from 7 to 15–20 minutes, 30 minutes recommended. Product page: games "each under 20 minutes". Visual/facial analysis was, by the company's public statement, withdrawn from the product (statement of January 2021; per SHRM, use ended in March 2020). Checked 4 August 2026.Source | No video interview and no face or voice analysis: only decisions inside the simulation are recorded. The client passes de-identified codes, so no name, gender or age reaches us — which also means we have no ATS connectors. |
| Criteria — Cognify | Class: mini-game battery. Standard version is three games and "takes just 10 minutes to complete, plus tutorials"; "a six-game version is also available on request". The Information Brief (© 2025) reports test–retest .81 (n = 280) and convergence with the vendor's RCAT at r = .301–.54 with about 500 participants per game. Product page lists English (US) by default plus Turkish, with uCognify offered as the language-independent version. Checked 4 August 2026.Source | Test–retest above 0.83 and Cronbach's α 0.69–0.77 — the α range is below the conventional 0.80 threshold, and we say so on the page rather than in a footnote. The client-facing reference library exists in Russian only. |
| SHL | Class: mixed. We did not encounter mini-games in the SHL materials we checked, and the word "games" does not appear in those factsheets. Verify Interactive G+ is an adaptive ability test with a drag-and-drop response format — 36 minutes, maximum 24 items (factsheet © 2018). Separately, job simulations in realistic interfaces: Contact Center Call Simulation, 2 tasks, 15 minutes average against a 20-minute limit (© 2020); Automata Pro, 2 tasks in an IDE, 46 minutes average against a 60-minute limit (© 2020). Checked 4 August 2026.Source | One 30–60 minute session with no adaptive item delivery: difficulty rises along the scenario, identically for everyone, so two candidates' sessions are directly comparable step by step. |
| Aon Assessment (formerly cut-e) | Class: modular mini-game battery. Vendor candidate page lists gridChallenge at 9 minutes and motionChallenge, digitChallenge, gapChallenge and switchChallenge at 6 minutes each, plus the gamified situational test chatAssess at 20 minutes. We found no fixed total duration on the vendor page we checked: the employer selects the modules. A complaint filed by the ACLU with the US Federal Trade Commission on 30 May 2024 concerns Aon's assessments; it is the complainant's position, and we have no information about an FTC decision. Checked 4 August 2026.Source | Duration is not configurable: 30–60 minutes in a single pass, with no modular assembly by the employer. We have no formal bias audit and no adverse impact ratio report of our own — stated in the closing section. |
| Test Partnership — MindmetriQ | Class: mini-game battery of six games — Net the Numbers 4 min, Number Racer 6, Link Swipe 6, Word Logic 6, Pipe Puzzle 4, Shape Spinner 4 — offered in three configurations of roughly 12, 18 and 35 minutes. The vendor publishes correlations with ICAR of 0.75–0.84 across test types, reaching 0.91 for verbal and 0.89 for numerical reasoning; we did not find sample sizes or a study design on that page. Checked 4 August 2026.Source | No separate verbal or numerical subtests: eight parameters are read from one scenario. We do not cover verbal and numerical ability and do not replace subject-matter knowledge testing. |
| Equalture | Class: mini-game battery assembled per role. Vendor FAQ: "a full assessment session usually takes between 15 and 45 minutes, depending on the number of games used for a role" and "each game itself lasts only a few minutes"; we did not find the total number of games in the library on the pages we checked. The Science page states development "according to the COTAN Review System" — a vendor statement of following a review framework; we did not find numeric validity coefficients on that page. Checked 4 August 2026.Source | The number of tasks is one and it is published: a single simulation of 30–60 minutes. The range comes from the person's own pace, not from how many modules were switched on. |
| Sova Assessment | Class: mixed. The Early Careers battery is stated at 25 minutes across seven components, of which the personality questionnaire and the situational judgement test are mandatory and the two gamified modules are marked optional. Separately, Sova Immerse is described as one continuous role simulation in three stages, "15–30 min"; we did not find links to validation studies on that page. Checked 4 August 2026.Source | There is no self-report questionnaire in the product at all: all eight parameters are derived from decisions inside the simulation. The trade-off is that we have no peer-reviewed publication of our own validation studies either. |
| NeuroFrame | This is our own product, so we do not rate it against the others — the verifiable characteristics are in the next column, and the section "When game-based assessment is the wrong tool" lists where we lose. | Class: one continuous strategy simulation, 30–60 minutes. Eight parameters — three cognitive, five personality — read from decisions, not from self-report. Percentiles against a comparison sample of 14,850 people in real jobs (500+ companies, 21 industries, 21 functions, 8 grades). CFI 0.96; Cronbach's α 0.69–0.77; test–retest above 0.83; R² 0.46 and AUC 0.77 against real KPIs and manager ratings on a predictive validation sample of 3,000+. R² 0.46 and AUC 0.77 were obtained on our own sample against our own criteria — they are not comparable with the coefficients of other instruments in this table and should not be read against them. Role library: 287 professions × 8 lifecycle stages = 2,296 records, each with its own human-readable rationale. |
Every product on this page is called by the name its rights holder sells it under: pymetrics — a product line within Harver; Arctic Shores — Arctic Shores Limited (United Kingdom); HireVue — HireVue, Inc.; Cognify and Emotify — products of Criteria Corp, which came to it together with Revelian; SHL — SHL Group Limited; Aon Assessment — formerly cut-e GmbH, part of Aon plc; MindmetriQ — a product of Test Partnership Limited; Sova — Sova Assessment Limited; Equalture — Equalture B.V. All names and trademarks belong to their respective owners and are used nominatively, to identify the product itself, with no logos and no trade dress.
Compared one by one
- One Long Simulation or a Battery of Mini-GamesTwo assessment architectures, checked 4 August 2026: what peer-reviewed work supports, what it refutes, and where a long session is the wrong tool.
- NeuroFrame and pymetrics (Harver)pymetrics (Harver): "12+" mini-games, 25 stated minutes; NeuroFrame: one 30–60 minute simulation. Every figure sourced, checked 4 August 2026.
- NeuroFrame and Arctic ShoresArctic Shores publishes up to nine short tasks; NeuroFrame runs one 30–60 minute simulation. Every vendor fact linked and dated, checked 4 August 2026.
- NeuroFrame and HireVueHireVue game-based assessments: format, duration, published validity and bias audits, each with a source link. Checked 4 August 2026.
- NeuroFrame and Criteria (Cognify)Cognify is three mini-games in a stated ten minutes; NeuroFrame is one 30–60 minute simulation. Every figure sourced, checked 4 August 2026.
- NeuroFrame and SHLVerify Interactive G+ is 36 minutes; NeuroFrame is one 30–60 minute session. Every SHL figure carries a source link. Checked 4 August 2026.
- NeuroFrame and Aon Assessment (cut-e)Aon Assessment (formerly cut-e): six gamified modules of 6–20 minutes vs one 30–60 minute simulation. Sourced facts, checked 4 August 2026.
What game-based assessment actually is
Game-based assessment is a selection procedure in which a candidate does not describe themselves on a form but performs interactive tasks, and the conclusion about their capabilities is derived from recorded behaviour: which decisions they made, how long each one took, and how their choices changed as load increased. The unit of measurement is an action, not an answer about oneself.
This is not the same thing as gamified recruiting. Points, badges and a leaderboard in a careers section raise engagement and measure nothing. In game-based assessment the mechanic is the carrier of the task: it fixes a set of rules inside which behaviour becomes comparable between people. Everything downstream — telemetry, the model that turns behaviour into scores, the norm group — is built the way any psychometric instrument is built.
The vocabulary across the category is not settled, and that is a practical problem for a buyer. Criteria describes its products literally as "short, fun, and engaging mini-games". On Arctic Shores pages as of 4 August 2026, two wordings sit side by side: the marketing line "tasks, not questions" and the glossary entry "our game-based assessment". SHL writes that "our immersive assessments upgrade button clicking with 'drag-and-drop' and gamified interactions" — and in the SHL factsheets we checked (© 2018–2024) we did not encounter the word "games". Three vendors, three vocabularies, three different things behind them.
So the useful question is not "is this a game?" but "what class of instrument is this?". That is what the rest of this page is organised around: classes first, brands second.
Four classes of game-based assessment tools
Products sold under the label "game-based assessment" fall into four classes: a battery of short mini-games, one long simulation, a mixed format, and a gamified self-report questionnaire. The class determines three things at once — what the instrument can observe, how long it occupies the candidate, and how far an individual task can be rehearsed in advance.
A mini-game battery is three to twelve independent tasks of one to six minutes each, assembled per role; the candidate finishes one mechanic and starts an unrelated one. One long simulation is a single continuous scenario in which earlier decisions change later conditions. A mixed format packages assessment components of different kinds into one flow — games plus a video interview, or a questionnaire plus a situational judgement test plus two gamified modules. A gamified questionnaire keeps the self-report as the unit of measurement and changes only the interface.
The boundary between the third and the fourth class is the slipperiest one, and it is where most buying mistakes happen. The reliable test is not the marketing word but the response-format line in the factsheet. SHL's Multitasking Ability factsheet (© 2018) describes a split-screen simulation of 20 minutes with 38 items — and records the question format as "Multiple Choice" (factsheet © 2018, checked 4 August 2026).
A class is not a quality rating. None of the four is shown by the current evidence to predict job performance better than the others: the systematic review of 34 studies by Ramos-Villagrasa, Fernández-del-Río and Castro (Frontiers in Psychology, 2022) concludes that game-related assessments do not offer enough advantage to be recommended in place of conventional methods, unless improving applicant reactions is itself the value. Choosing a class is a choice about what you get to observe and what you spend, not a shortcut to accuracy.
Class 1. A battery of short mini-games
A battery of short mini-games is the most common format among the products we checked on 4 August 2026. By the vendors' own published pages: pymetrics inside Harver states "12+ interactive, gamified experiences" and "25 Minutes to complete"; Criteria Cognify's standard version is three games and "takes just 10 minutes to complete, plus tutorials", with a six-game version available on request; Test Partnership MindmetriQ is six games of four to six minutes in configurations of roughly 12, 18 and 35 minutes; Equalture's FAQ states "a full assessment session usually takes between 15 and 45 minutes, depending on the number of games used for a role"; Aon Assessment lists gridChallenge at 9 minutes and motionChallenge, digitChallenge, gapChallenge and switchChallenge at 6 minutes each, with the combination chosen by the employer.
The class measures one thing very well: isolated cognitive functions with a long research history outside hiring. The peer-reviewed FAccT 2021 audit, written jointly by Northeastern researchers and pymetrics staff, names the twelve games and their underlying paradigms — Balloon Analogue Risk Task, Trust Game, Go/No-Go, digit span, Iowa Gambling Task, Flanker, Tower. Each of those paradigms was studied for decades before anyone used it for selection, and each has a literature describing what it reacts to.
Modularity is the second real strength. A short battery can be reconfigured per role without rebuilding the instrument, which is why the vendors in this class whose pages we checked on 4 August 2026 publish a range rather than a single number. Short sessions also cost the funnel less: assessment length is one of the few levers a recruiter controls directly.
What the class cannot do follows from the same design. Each task is self-contained, so nothing carries over: a decision made in minute two has no consequence in minute nine. There is no work context around the mechanic, and behaviour is strongly situation-dependent — in assessment centre research, variance in ratings is explained by the specific exercise at least as much as by the competency being rated (Lievens and colleagues, 2006). A battery also observes the candidate only while fresh; there is no long stretch on which sustained attention could visibly decline.
For the candidate this class is the least demanding: short, mobile-friendly, easy to fit into a lunch break. For the employer this class is simpler to deploy, and around named tasks — as the search-demand section below shows — a preparation industry exists.
Class 2. One long simulation
A single long simulation is one continuous scenario in which the candidate's earlier decisions change the conditions of later ones. Among the products we checked on 4 August 2026, we did not find a vendor building high-volume screening on one uninterrupted strategic simulation of 30 minutes or more; the class exists on the market, but mostly in shorter or narrower forms.
The closest published products are three. Sova Immerse is described by its vendor as a single continuous role simulation in three stages — manager briefing, task simulation, manager debrief — of "15–30 min". CapsimInbox is one inbox simulation: "typical inbox simulations are completed in just 15-60 minutes", and Capsim lists Hiring & Selection among its use cases. Owiwi is a single narrative adventure, "Isles of the Shroud", built as eight islands; we found no duration — a full-text search of its 30-page technical manual returns no occurrences of "minute" or "duration" (checked 4 August 2026) — so the format is comparable but the length is not.
What a continuous scenario can do that a battery cannot is measure complex problem solving: performance in dynamic micro-worlds with delayed consequences overlaps with general cognitive ability by roughly 18% of variance (r ≈ .43; meta-analysis of 47 studies, Stadler et al., 2015), which means it also contains something classical tests do not capture. A long session also lets an employer watch behaviour both fresh and tired: the decline of sustained attention with time on task is one of the most reproducible effects in cognitive psychology at the group level. And engagement builds through a session rather than being set in the first minute (Genc et al., 2026).
What this class cannot do is claim higher predictive validity, and the honest reading of the data points the other way. On the re-estimated meta-analytic figures of Sackett, Zhang, Berry and Lievens (2022), work samples sit at .33 and assessment centres at .29 — below structured interviews at .42 and job knowledge tests at .40. Nor should a long session be sold as measuring accumulated fatigue as an individual trait: a change score between the start and the end of a session has poor test–retest reliability, so it belongs in the observation, not in the scale.
For the candidate this class costs more time and more attention, and it does not fit a five-minute gap between meetings. For the employer it is harder to norm — one long instrument produces one sample, not six — and it does not slice into modules you can drop for a junior role.
Class 3. Mixed formats
A mixed format packages components of different kinds into a single candidate flow, and the games are one component among several rather than the product itself. Three of the vendors in our table are built this way: HireVue, SHL and Sova.
At HireVue the game block is a module inside the assessment platform. The vendor's ebook describes five pre-built cognitive packages and names three of them — Cognition Short, Cognition Standard, Cognition Comprehensive — and states that the game-based assessments "take under 15 minutes to complete"; we did not find a published duration for each individual package. The candidate-facing page says each game takes approximately 3 minutes and advises setting aside 30 minutes; the current product page says the games run "each under 20 minutes". The typical deployment described in HireVue's own assessment-science whitepaper combines short games with a video interview in one flow.
SHL is an example of a mixed portfolio in which we found no mini-games. Verify Interactive is a classical adaptive ability test in which only the response format changed: G+ is 36 minutes and a maximum of 24 items by the © 2018 factsheet. Separately, and built on a different principle, are job simulations in realistic interfaces — Contact Center Call Simulation, 2 tasks, 15 minutes average against a 20-minute limit (© 2020); Automata Pro, 2 tasks, 46 minutes average against a 60-minute limit in an IDE (© 2020). In these documents (© 2018 and © 2020) we did not find the word "games" — checked 4 August 2026.
Sova shows how much the proportion matters. Its Early Careers battery is stated at 25 minutes across seven components, of which the personality questionnaire and the situational judgement test are mandatory and the two gamified modules — Gamified Numerical Challenge and Gamified Pattern Challenge — are marked optional (checked 4 August 2026).
The practical consequence is one question to ask any vendor in this class: which share of the session is game-based, and which component actually produces the score that gets used. The second consequence is arithmetic — packages add up, and a mixed stack of a 36-minute ability test plus a 16-minute questionnaire is a different candidate experience from a 10-minute battery.
Class 4. A gamified self-report questionnaire
A gamified self-report questionnaire keeps the questionnaire as the unit of measurement and changes the shell around it: animation, timers, a scenario framing, a progress bar. The candidate is still telling the system about themselves; the system is not watching them solve anything.
The reliable way to tell this class from a simulation is the response-format field in the vendor's own factsheet, not the adjective on the product page. SHL's Multitasking Ability (© 2018) is described as a split-screen simulation, 20 minutes, 38 items — with question format "Multiple Choice". SHL's Professional 8.0 Job-Focused Assessment (© 2024) is 16 minutes and up to 76 forced-choice items, and the Global Skills Assessment factsheet (v1.0, 21 May 2024) states that GSA "is used in all 8.0 Job Focused Assessments (JFAs)" — 15 minutes, 76 forced-choice triplets. SHL calls these products skills assessments; the response-format field in the same factsheets records forced choice. Both statements come from the vendor's own documents and are reproduced here without interpretation.
The class has genuine advantages, and they are not small. Self-report questionnaires are cheap to administer, fast to complete, easy to explain to a candidate and to a works council, straightforward to localise into many languages, and backed by decades of construct research and large norm groups. For many roles that is exactly the right trade.
Its limit is the one it has always had: the candidate controls the answer. Gamifying the format helps but does not solve it — in a controlled experiment (N = 171) a gamified measure of a personality trait showed a significantly smaller deliberate-distortion effect than the conventional questionnaire, and distortion was still possible. "Reduces faking" and "prevents faking" are different claims, and only the first one is supported.
The boundary of the class is worth naming too, because not everything short and online belongs in this category. The McQuaig Word Survey takes approximately 20 minutes by the vendor's own page, and on that page we found no mention of game mechanics — the vendor describes it as a conventional instrument (checked 4 August 2026).
Why assessment sessions kept getting shorter
The market moved toward shorter sessions, not longer ones. Criteria now sells Cognify as three games and "just 10 minutes to complete, plus tutorials", with the six-game version available on request. HireVue's ebook says its game-based assessments "take under 15 minutes to complete", and the vendor's own whitepaper puts a full battery at 6–15 minutes. Test Partnership's home page offers a "Complete cognitive profile in 14 mins". Arctic Shores states "know true potential, in 45 minutes" on its own How it works page, while an employer's public candidate briefing we checked on 4 August 2026 asks candidates to set aside 15–20 minutes, because the number of tasks is configured by the employer.
The reason is funnel economics, and it is not mysterious. Assessment length is one of the few variables a talent team controls directly, it sits at the top of the funnel where volume is largest, and it is the variable candidates complain about. Vendors compete on it explicitly: completion-rate claims appear on the marketing pages of assessment platforms — Traitify (Crosschq), for example, states a 95% completion rate on its site (traitify.com, checked 4 August 2026). That is the company's own statement, and we did not find a description of the methodology or the sample on that page.
What gets traded away is anything that only becomes visible over a long stretch: how decision quality holds up as load accumulates, whether a person re-plans after a setback, what they do in the fourth consecutive difficult minute. A ten-minute battery cannot show that, not because it is built badly but because there is no long stretch in it.
It is worth saying plainly that shortening was a rational response to a real problem, not a decline in standards. The published evidence does not show that longer sessions predict job performance better — on Sackett et al. (2022) the higher-fidelity, longer methods sit below structured interviews. Session length is an argument about what you observe and what it costs the funnel, and it should be made in those terms.
The test-prep industry as a measurable property of the format
Around batteries of named mini-games there is a preparation industry, and its size is measurable rather than a matter of opinion. Per Semrush (UK database, pulled 4 August 2026), search demand for walkthroughs of individual Arctic Shores tasks runs at roughly 550 queries a month in total — balance 170, order 110, predict 110, lock 90, ticket 70. "Arctic shores practice test free" adds 390 a month, and related preparation queries for that one vendor total more than 1,400 a month. "Pymetrics games answers" runs at 110 a month in the US database of the same tool.
The supply side matches the demand. In Google's US results for the head query "game based assessment" on 4 August 2026 we saw the first organic position held by gameassessmentprep.com — a preparation site. Results are personalised and change over time, so this is an observation from our own check, not a permanent property of the market.
What this proves is narrow and specific: a discrete, named, repeatedly used task can be found, described and rehearsed. That is a structural property of the class, not a defect of any one product. A mechanic that has a name and a stable rule set is a mechanic somebody can write a guide about.
What it does not prove is that preparation changes the score, and the honest treatment requires saying so. Criteria itself reports in its Cognify Information Brief (© 2025) a "weak to moderate positive relationship" between weekly gaming hours and performance on Grid Lock, and characterises the strength of that link as minimal and comparable to the effect of prior experience with conventional psychometric tests. In other words the vendor reports trainability itself and characterises the effect as minimal.
The practical consequence for an employer is a short list of questions: how often are items rotated, what is the item-security policy, is there a published measure of practice effects, and does the score used in the decision come from a mechanic with a public walkthrough. The consequence for a candidate is that preparation for this class exists and is legal — and that the Russian-language market behaves differently: "как пройти тест при приеме на работу" runs at about 10 queries a month in the same tool's Russian database (pulled 4 August 2026), an order of magnitude below the English demand, so in Russia the sharper question is completion, not rehearsal.
What is actually known about validity
As of 4 August 2026 we found no meta-analysis of the predictive validity of game-based assessment against job performance. That is the honest state of the field as we were able to establish it, and on our checking it holds for the instruments in the table below, ours included. The systematic review of 34 studies by Ramos-Villagrasa, Fernández-del-Río and Castro (Frontiers in Psychology, 2022) concludes that game-related assessments do not offer enough advantage to be recommended in place of conventional selection methods, unless improving applicant reactions is itself considered added value — and recommends using only instruments designed for selection from the start and grounded in psychological theory.
What has been established is narrower and still useful. Meta-analytic work links game-based measures to conventional ones — Bipp and colleagues (2024) report a corrected r = .45 with cognitive ability tests, and Fadillah and colleagues (2025) report r = .52 with personality questionnaires. The same literature finds that the game format does not produce greater adverse impact than conventional formats, and that candidates react to it more positively. That is a real result, and it is the result the review says you should buy the format for.
For comparison, the current meta-analytic estimates for selection methods generally come from Sackett, Zhang, Berry and Lievens (2022): structured interview .42, job knowledge tests .40, work samples .33, general cognitive ability .31, assessment centres .29, situational judgement tests .26, conscientiousness .19. The older Schmidt and Hunter (1998) figures — GMA .51, work samples .54 — were inflated by systematic over-correction for range restriction, on the same authors' re-analysis, and should not be quoted as current.
What vendors do publish is mostly convergent validity: evidence that the game measures roughly the same thing as an established test. Test Partnership publishes correlations of MindmetriQ with ICAR: on the vendor page as of 4 August 2026 they run at 0.75–0.84 across test types and reach 0.91 for verbal and 0.89 for numerical reasoning. We did not find sample sizes or a study design on that page. HireVue's 2021 whitepaper reports cross-validated multiple R of .51–.67 against ICAR on modelling samples of 364–647, and a peer-reviewed Frontiers paper (2023, three of four authors affiliated with HireVue) reports r = 0.5 with ICAR and test–retest r = 0.68 on 102 participants. Criteria's Cognify brief reports convergence with its own RCAT at r = .301–.54 with about 500 participants per game and test–retest .81 (n = 280). Aon's 2019 presentation reports r = .71 between gridChallenge and its predecessor G.A.M.E. on N = 306 recruited via Mechanical Turk.
Convergent validity answers the question "are we measuring roughly the same construct?". It does not answer "does this predict performance in the job?". Those are different questions with different study designs, and a buyer comparing vendors should ask which of the two a given number belongs to — and on what sample, in what year, published where.
What the EU AI Act changed on 2 August 2026
From 2 August 2026 the main body of the EU AI Act applies, and systems used to evaluate candidates in recruitment are classified as high-risk under Annex III, point 4(a). One thing changed shape shortly before that date: the obligations attaching to Annex III high-risk systems were postponed from 2 August 2026 to 2 December 2027 by Regulation (EU) 2026/1744, published in the Official Journal on 24 July 2026. The classification did not move; the compliance deadline did.
Two things are in force regardless of that postponement. The transparency obligations under Article 50 apply from 2 August 2026. And since 2 February 2025 the Act prohibits AI systems that infer a person's emotions in the workplace — a prohibition that was not affected by the delay. The prohibition is framed around inferring emotions, and whether a particular product falls under it is a matter for its provider and the supervisory authority; we make no such determination about any product. The primary source is linked below.
The EU AI Act is not the only instrument that applies here, and for a hiring decision it may not be the binding one. Article 22 GDPR already prohibits, by default, decisions based solely on automated processing that significantly affect a person, and requires at minimum the right to human intervention, to express a point of view and to contest the decision. In SCHUFA (C-634/21, 7 December 2023) the Court of Justice held that Article 22 obligations can also fall on the provider of a score, where the recipient draws strongly on it. In Dun & Bradstreet Austria (C-203/22, 27 February 2025) the Court held that trade secrecy does not excuse a controller from giving a person an explanation of the logic of an automated decision sufficient to contest it.
In the UAE there is no dedicated law on AI in hiring; the binding norm is Article 18 of the federal Personal Data Protection Law, which gives a person the right to object to a decision made by automated processing, including profiling. Inside the DIFC, Regulation 10 — in force since 1 September 2023 — requires disclosure of the use of an autonomous system, documentation of bias-detection and human-intervention mechanisms, and, for high-risk processing, certification and a designated responsible person.
Across all of these regimes the requirement converges on the same three things, and none of them is "do not use AI": explainability of how a conclusion was reached, independent checking for differential impact, and a real human able to intervene in the decision. Nothing on this page is legal advice; the primary sources are linked below and should be read against the date on which you intend to deploy.
Who publishes bias audits, and what those audits mean
Under New York City Local Law 144 the duty to commission and publish a bias audit falls on the employer or employment agency, not on the assessment vendor. A vendor that publishes an audit is therefore sending a voluntary signal of maturity, not complying with a legal obligation of its own. The law also defines independence strictly: an auditor tied to the tool's development, to an employment relationship or to a financial interest does not count as independent, which rules out internal self-checks.
Several vendors in this table do publish. Harver publishes a full independent audit by BABL AI dated 17 July 2025 covering the pymetrics soft-skills platform, with all disclosed impact ratios between 0.914 and 1.000; the same report states that testing was performed by Harver in June 2025 and that BABL AI verified the preparation of the statements, that 283,256 candidates were included in the gender analysis, and that 331,177 were excluded because at least one demographic field was missing. A DCI Consulting Group report on HireVue, published by an employer in August 2023, contains a separate section for the game tool "Think – Shapedance, Numerosity" with impact ratios of 0.86–1.00 (checked 4 August 2026). For Arctic Shores the audit we found was published not by the vendor but by a client employer, FDM Group.
Published audits are worth reading with their own limits in view. The FAccT 2025 study "Auditing the Audits" analysed 44 published reports containing 116 audits between July 2023 and early November 2024, covering roughly 2% of Fortune 500 companies; 53% of the published audits contain at least one impact ratio below 0.8, and 83% acknowledge missing demographic data. The same authors documented repeated identical results across reports and stated explicitly that they could not establish the cause, noting that legally permitted pooling of several employers' data by one auditor is a possible explanation.
The four-fifths rule itself is not a certificate. Under the 1978 Uniform Guidelines a ratio above 0.8 is generally not treated as evidence of adverse impact, but smaller differences can still constitute it where they are statistically and practically significant. "Passed an audit" and "has no differential impact" are different statements.
For a buyer this turns into four concrete questions: who conducted the audit and what disqualifies them from independence, what period and what sample it covers, how many candidates were excluded for missing demographics, and whether the module actually used in your decision appears in the report by name. A vendor audit that does not name the specific instrument does not tell you about the specific instrument.
Nine questions that tell the classes apart
Marketing vocabulary in this category does not separate the classes; a short list of factual questions does. Every one of these can be answered from a factsheet or a product page, and a vendor unwilling to answer any of them has told you something useful.
On format: (1) How many separate tasks are there, and are they named? (2) What is the response format for each — an action inside a scenario, or a choice among options? (3) Do earlier decisions change later conditions, or is each task self-contained? (4) What is the total duration as a range, and what determines where inside the range a given candidate lands?
On evidence: (5) Which published numbers are convergent validity against another test, and which are criterion validity against job outcomes? (6) What was the sample size and the year, and where was it published? (7) What is the norm group — its size, its composition, and whether it is a general norm or a model trained on this employer's current high performers.
On governance: (8) Which module produces the score that enters the decision, and does that module appear by name in any published bias audit? (9) What personal data does the vendor receive, and what explanation can the candidate be given if they contest the outcome — in a form that satisfies Article 22 GDPR and, from December 2027, the high-risk regime of the EU AI Act.
The answers place a product in one of the four classes within about ten minutes, and they also reveal the two most common mismatches: a battery bought in the belief that it observes sustained performance, and a self-report bought in the belief that it observes behaviour.
Where NeuroFrame sits on this map
NeuroFrame belongs to the second class: one continuous strategy simulation of 30–60 minutes, from which eight parameters are read — three cognitive (Mental Efficiency, Learning Agility, Progress Monitoring) and five personality (Openness to New, Risk Appetite, Result Focus, Agreeableness, Conscientiousness). Scores are percentiles against a comparison sample of 14,850 people in real jobs across 500+ companies, 21 industries, 21 functions and 8 grades. There is no mini-game battery and no self-report questionnaire in the product.
The role benchmark is built from the outside rather than from the client's staff. The reference library holds 287 professions from the Russian Ministry of Labour occupational register across 34 functional areas, crossed with 8 stages of the Adizes corporate lifecycle — 2,296 records. The consequence is practical: the method works from day one, without needing dozens of current employees in the role, and it does not structurally encode the existing composition of a workforce. The same data show why the stage matters: for 99% of the 287 professions, the requirement ranges at the Infancy and Bureaucracy stages overlap by less than half, with a median range overlap of 0.40, and of the 153 outright non-overlaps between extreme stages, 152 fall on risk appetite.
Every one of the 2,296 cells carries its own human-readable rationale — 2,296 unique texts, 1.07 million characters, 464 characters per cell on average, and 32% of them name the rule that produced the range. In 303 of the 2,296 cells (13%) the benchmark permits a high risk appetite alongside a low floor on thinking or monitoring, which is how the method can say "fits the role on every parameter and still warrants attention" — something a single aggregate fit percentage cannot express by construction.
De-identification is architectural, not a compliance line. The client receives codes and distributes them internally; NeuroFrame receives no name, no gender and no age. A model that never sees a protected attribute cannot use it directly. That closes direct use; it does not by itself rule out indirect effects — we have no formal bias audit, and we say so in the closing section. The trade-off is stated there too: it is also why there are no ATS connectors.
The psychometrics, with their caveats: CFI 0.96 for a two-domain confirmatory model, Cronbach's α 0.69–0.77, test–retest above 0.83, and R² 0.46 with AUC 0.77 against real KPIs and manager ratings on a predictive validation sample of 3,000+. Two of those numbers deserve to be read carefully. The α range sits below the conventional 0.80 threshold — we say so in the closing section rather than in a footnote. And our own R² was obtained on our own sample against our own criteria: it is not comparable with meta-analytic coefficients from other instruments, and we do not put it in the same row as them.
Conflict of interest disclosure
This comparison is published by NeuroFrame — that is, by an interested party. Every fact about another company's product therefore carries a link to its primary source, and we do not score anyone at all.
That constraint shaped the whole page. There are no evaluative adjectives about anyone else's product on it: not "weak", not "shallow", not "low validity". There is no ranking, no star rating and no summary verdict. Every number about another vendor appears alongside the vendor's own document and the date on which we checked it, 4 August 2026. Where our search did not find something, we say that our search did not find it — which is a statement about our search, not about what exists.
There are three things we deliberately do not do. We do not publish competitors' prices from third-party procurement catalogues, because those figures age and diverge from the vendor's own terms. We do not describe any vendor as discriminating, biased or non-compliant: where complaints have been filed with regulators, we state that a complaint was filed and that we have no information about a regulator's decision. And we do not attach Review, AggregateRating or Product markup to anyone else's product.
The page is also checkable against us. Every vendor claim here carries a link to its primary source and the date we checked it — 4 August 2026 — and every NeuroFrame number comes from a single internal source of truth rather than from a marketing draft. If you find a statement that has gone stale or that we got wrong, write to us: a correction with a date is worth more to this page than the original sentence was.
When game-based assessment is the wrong tool — and when we are
Game-based assessment is the wrong tool whenever the decision turns on a checkable skill. If the question is whether someone can write the code, close the books or hold a conversation in the language, use a job knowledge test or a work sample: on Sackett et al. (2022) those sit at .40 and .33, and they answer the actual question directly. Behavioural instruments, ours included, are not built to tell you whether a person knows a syntax or a standard.
It is also the wrong tool at low volume. If you hire five people a year for one role, a structured interview — .42, the highest coefficient on that list — is cheaper, faster and easier to defend than deploying any assessment platform. Assessment platforms earn their cost in the wide part of the funnel, where the number of candidates makes consistency worth buying.
And it is the wrong tool where the candidate cannot fairly take it. Any timed, real-time format disadvantages people who need adjustments it does not offer. If you cannot provide an alternative route for those candidates, do not make the assessment the gate.
Where we specifically are the wrong choice, stated without softening. We have no peer-reviewed publication of our own validation studies. We have no formal bias audit and no adverse impact ratio report — and until we do, we make no claim of zero bias. Our internal consistency is 0.69–0.77, below the conventional 0.80 threshold. Our comparison sample is 14,850 people, smaller than instruments with forty-year histories. Our compliance perimeter is Russian: 152-FZ is closed architecturally because we receive no personal data at all, but we do not yet have a GDPR, EU AI Act or NYC Local Law 144 package. We have no ATS connectors, which is a direct consequence of the de-identified-code design. We measure with one task, so we do not cover verbal and numerical abilities and do not replace subject-matter testing. Our adjustments policy for candidates with disabilities is not yet formalised. 83% of the library records are specialist-level roles, with top management covered by nine professions. And the client-facing reference library exists only in Russian.
If any of that is disqualifying for your case, it should disqualify us, and we would rather say it before a pilot than after one. If a vendor in the table above fits your constraints better, the links are there and they lead to that vendor's own pages, not to ours.
Questions
- What is game-based assessment, in plain terms?
- It is a selection procedure in which the candidate performs interactive tasks instead of describing themselves on a form, and conclusions are drawn from recorded behaviour: which decisions were made, how long each took, and how choices changed as load increased. It differs from gamified recruiting in that the mechanic carries the task rather than decorating it — it fixes rules inside which behaviour becomes comparable across people.
- How is game-based assessment different from a gamified questionnaire?
- By the unit of measurement. Game-based assessment measures an action inside a task; a gamified questionnaire still measures the person's statement about themselves, in a nicer shell. The test is not the marketing word but the response-format line in the factsheet: SHL's Multitasking Ability factsheet (© 2018) describes a 20-minute split-screen simulation and records the question format as "Multiple Choice".
- Can candidates train for a game-based assessment?
- A measurable preparation industry exists around batteries of named mini-games: walkthroughs of individual Arctic Shores tasks draw roughly 550 UK queries a month, "arctic shores practice test free" adds 390, "pymetrics games answers" runs at 110 a month in the US (per Semrush, pulled 4 August 2026). That proves a discrete named task can be found and studied. Whether practice moves the score is a separate question: Criteria itself reports a "weak to moderate positive relationship" between gaming hours and Grid Lock performance and calls the effect minimal.
- How long does a game-based assessment take?
- By the vendors' published data on 4 August 2026: Criteria Cognify 10 minutes plus tutorials; HireVue's game-based assessments "under 15 minutes" by the vendor's ebook; Test Partnership MindmetriQ roughly 12, 18 or 35 minutes; pymetrics inside Harver 25 minutes; Equalture 15–45 minutes; Sova Early Careers 25 minutes; Arctic Shores states "know true potential, in 45 minutes" on the vendor page, with 15–20 minutes in an employer's candidate briefing. NeuroFrame is 30–60 minutes in a single pass.
- Is game-based assessment more accurate than conventional tests and questionnaires?
- Not on the evidence available in 2026. The systematic review of 34 studies (Ramos-Villagrasa et al., 2022) concludes that game-related assessments do not offer enough advantage to be recommended in place of conventional methods, unless improved applicant reactions are themselves the added value. As of 4 August 2026 we found no meta-analysis of predictive validity against job performance. What is established: the format does not produce greater adverse impact, and candidates react to it markedly better.
- What did the EU AI Act change on 2 August 2026?
- From 2 August 2026 the main body of the Regulation applies, and systems evaluating candidates in recruitment are high-risk under Annex III, point 4(a). The obligations attaching to high-risk systems, however, were postponed to 2 December 2027 by Regulation (EU) 2026/1744 of 8 July 2026, published in the Official Journal on 24 July 2026. Independent of that postponement, the Article 50 transparency obligations apply from 2 August 2026 and the prohibition on systems inferring emotions in the workplace has applied since 2 February 2025. This is a description of the requirement, not legal advice.
- Is a vendor required to publish a bias audit?
- Under NYC Local Law 144 the duty to commission and publish the audit sits with the employer or employment agency, not with the vendor. A vendor publishing an audit is sending a voluntary maturity signal. The law defines independence strictly: an auditor tied to the tool's development, to an employment relationship or to a financial interest does not qualify.
- Does game-based assessment replace the interview?
- No. On the current meta-analytic estimates (Sackett et al., 2022) the structured interview remains the highest-coefficient method at .42, above job knowledge tests (.40), work samples (.33) and assessment centres (.29). The sensible role for game-based assessment is a consistent, comparable read at the wide end of the funnel, after which the interview starts from specific observations rather than from scratch.
- Which vendor belongs to which class?
- By the vendors' published data on 4 August 2026: mini-game batteries — pymetrics (Harver), Criteria Cognify, Test Partnership MindmetriQ, Equalture, Aon Assessment's Challenge line, and Arctic Shores, which calls its items tasks; mixed formats — HireVue, SHL and Sova; one continuous simulation — Sova Immerse (15–30 minutes), CapsimInbox (15–60 minutes) and Owiwi (we found no duration in the vendor's public materials). The full table with links is above.
Sources
Every link was opened and checked on the date shown above.
- Harver — Gamified Assessments (pymetrics) — Harver
- Arctic Shores — How it works — Arctic Shores Limited
- Arctic Shores — Glossary: game-based assessments — Arctic Shores Limited
- HireVue — Game-Based Assessments — HireVue, Inc.
- HireVue — How to prepare for your HireVue assessment — HireVue, Inc.
- HireVue — How Game-Based Assessments Uncover Top Talent (ebook): five pre-built cognitive packages incl. Cognition Short / Standard / Comprehensive; "take under 15 minutes to complete"; 3 minutes per game — HireVue, Inc.
- HireVue — The Next Generation of Assessments whitepaper (October 2021): a complete game-based battery "takes only 6-15 minutes to complete" — HireVue, Inc.
- Test Partnership — home page ("Complete cognitive profile in 14 mins") — Test Partnership Limited
- Criteria — Cognify — Criteria Corp
- Criteria — Cognify Information Brief (© 2025) — Criteria Corp
- Criteria — Assessments catalogue ("short, fun, and engaging mini-games") — Criteria Corp
- SHL — Verify Interactive G+ factsheet (© 2018) — SHL Group Limited
- SHL — Multitasking Ability factsheet (© 2018): 20 minutes, 38 questions, Question Format "Multiple Choice" — SHL Group Limited
- SHL — Professional 8.0 Job-Focused Assessment factsheet (© 2024): 16 minutes, max 76 questions, Question Format "Forced-Choice" — SHL Group Limited
- SHL — Global Skills Assessment factsheet (v1.0, 21 May 2024): 15 minutes, 76 triplets, Forced Choice, "used in all 8.0 Job Focused Assessments (JFAs)" — SHL Group Limited
- SHL — Contact Center Call Simulation factsheet (© 2020): 2 questions, 15 minutes average, 20 minutes allowed — SHL Group Limited
- SHL — Automata Pro factsheet (© 2020): 2 questions, 46 minutes average, 60 minutes allowed, IDE simulation — SHL Group Limited
- SHL — Cognitive Assessments ("Our immersive assessments upgrade button clicking with 'drag-and-drop' and gamified interactions") — SHL Group Limited
- Aon — Prepare for your online assessment (Challenge tests, chatAssess) — Aon plc
- Siemsen, smartPredict — Development Insights and Study Results of Aon's Gamified Assessment Series (GBA Workshop, Minneapolis, August 2019): G.A.M.E./gridChallenge equivalence r = .71, 306 participants with usable data, MTurk sample — Aon's Assessment Solutions / GBA Workshop
- Test Partnership — Gamified Assessment (MindmetriQ) — Test Partnership Limited
- Equalture — FAQ (session length, games per role) — Equalture
- Equalture — Science (COTAN Review System statement) — Equalture
- Sova Assessment — Early Careers (25 minutes, seven components) — Sova Assessment Limited
- Sova Immerse — 15–30 min role simulation in three stages — Sova Assessment Limited
- CapsimInbox — Inbox Simulations (15–60 minutes; Hiring & Selection) — Capsim Management Simulations
- Owiwi — Product ("Isles of the Shroud" narrative, eight islands) — Owiwi
- Owiwi — Soft Skill technical manual (norm sample 5,371; no duration found in the manual) — Owiwi
- McQuaig Word Survey — approximately 20 minutes, no game mechanics — The McQuaig Institute
- Wilson et al., Building and Auditing Fair Algorithms (ACM FAccT 2021) — twelve games, per-client models — ACM FAccT 2021
- Ramos-Villagrasa, Fernández-del-Río & Castro — Game-related assessments for personnel selection: a systematic review (Frontiers in Psychology, 2022) — Frontiers in Psychology / PMC
- Frontiers in Psychology (2023) — HireVue game-based cognitive assessment: convergent validity r = 0.5 with ICAR, test–retest r = 0.68 — Frontiers in Psychology / PMC
- HireVue's Assessment Science whitepaper (October 2021) — Multiple R .51–.67 against ICAR — HireVue, Inc.
- SHRM — HireVue discontinues facial analysis screening (announced January 2021; feature discontinued March 2020) — SHRM
- Criteria Corp — press release on the acquisition of Revelian (game-based assessments, Emotify) — Criteria Corp
- Southeastern — Arctic Shores test information sheet for candidates ("set aside anywhere between 15-20 minutes"; time varies with the number of tasks) — Southeastern (Go-Ahead)
- Bipp, Wee, Walczok & Hansal (2024) — The Relationship Between Game-Related Assessment and Traditional Measures of Cognitive Ability: A Meta-Analysis (observed r = .30, corrected r = .45), Journal of Intelligence 12(12), 129 — Journal of Intelligence / PMC
- Fadillah, Hidayat & Santoso (2025) — Convergent Validity of Game-Based Assessment: A Meta-Analysis (r = .516 against self-report personality measures; 18 studies), International Journal of Serious Games 12(4) — International Journal of Serious Games
- Stadler, Becker, Gödker, Leutner & Greiff (2015) — Complex problem solving and intelligence: A meta-analysis (47 studies, average effect size .43), Intelligence 53, 92–101 — Intelligence (Elsevier)
- Sackett, Zhang, Berry & Lievens (2022) — Revisiting meta-analytic estimates of validity in personnel selection — Journal of Applied Psychology 107(11), 2040–2068
- Sackett et al. (2023) — follow-up on revised validity estimates — Industrial and Organizational Psychology
- Vigilance decrement and time on task — review — PMC
- Genc et al. (2026) — continuous measurement of flow during a gaming session — Psychophysiology
- Controlled experiment (N = 171) — gamified measure shows smaller deliberate-distortion effect than a conventional questionnaire — Journal of Business and Psychology
- Assessment centre research — exercise variance versus dimension variance — PubMed
- Regulation (EU) 2026/1744 — postponement of AI Act high-risk obligations to 2 December 2027 (OJ, 24 July 2026) — EUR-Lex
- EU AI Act — Annex III, point 4(a): employment and worker management as high-risk — artificialintelligenceact.eu
- Prohibition on emotion recognition in the workplace under the EU AI Act (in force 2 February 2025) — Future of Privacy Forum
- GDPR Article 22 — automated individual decision-making, including profiling — gdpr-info.eu
- CJEU SCHUFA (C-634/21, 7 December 2023) — automated decision-making rulings, key takeaways — IAPP
- CJEU Dun & Bradstreet Austria (C-203/22, 27 February 2025) — trade secrecy does not bar explanation — Court of Justice of the European Union
- NYC DCWP final rules on Automated Employment Decision Tools (Local Law 144) — NYC Department of Consumer and Worker Protection
- BABL AI — pymetrics Soft Skills Platform 2025 Bias Audit (17 July 2025) — BABL AI Inc. / Harver
- DCI Consulting — HireVue bias audit summary (July 2023), including "Think – Shapedance, Numerosity" — DCI Consulting Group
- FDM Group — NYC applicant bias audit (Arctic Shores assessment, published by the employer) — FDM Group
- Gerchick et al., Auditing the Audits (ACM FAccT 2025) — 44 reports, 116 audits, 53% with an impact ratio below 0.8 — ACM FAccT 2025
- Uniform Guidelines on Employee Selection Procedures, 29 CFR 1607.4 — the four-fifths rule — Cornell Legal Information Institute
- AI in the UAE — regulatory landscape; Article 18 of the federal PDPL — Latham & Watkins
- DIFC Regulation 10 — personal data processed through autonomous and semi-autonomous systems — Mayer Brown
- ACLU complaint to the US Federal Trade Commission regarding Aon Consulting, Inc. (filed 30 May 2024) — American Civil Liberties Union
- Traitify (Crosschq) — company completion-rate claim: "a 95% completion rate" — Traitify / Crosschq