Comparison

NeuroFrame and pymetrics (Harver)

pymetrics is a battery of short mini-games; NeuroFrame is a single continuous simulation. Almost every other difference follows from that one. This page sets out what each side publishes, with a link next to every number and the date we checked it. It is published by NeuroFrame — an interested party — so we have tried to do without evaluations of the other product: sourced facts only, plus a section on where we are the weaker choice. If you find an evaluative phrase on this page, tell us and we will remove it.

Updated:

Side by side

AttributepymetricsNeuroFrame
Instrument classA battery of short mini-games; the vendor calls them "12+ interactive, gamified experiences" (browser and mobile app, per the FAccT 2021 paper, product as of summer 2020)SourceOne continuous strategic simulation — a single task rather than a battery
Number of tasks"12+" on the vendor's page, which does not state an exact current number; the ACM FAccT 2021 paper describes "a core set of twelve games" (product as of summer 2020)SourceOne session; 8 measured parameters — 3 cognitive and 5 personality
Duration"25 Minutes to complete", per the vendor. We found no vendor documentation with a per-game breakdown; the "one to three minutes per game" figure comes from the commercial prep site strategycase.comSource30–60 minutes in one continuous session
Normative modelA model trained per client: the in group is typically 50–100 high-performing employees of that client, with adverse-impact screening on a separate bias group typically exceeding 10,000 users (product as of summer 2020)SourceAn external reference: 287 professions × 8 corporate lifecycle stages = 2,296 benchmark cells, built from the Russian Ministry of Labour occupational register and the Adizes lifecycle model; comparison sample of 14,850 people
Published bias audit (NYC LL 144)BABL AI Inc., report dated 17 July 2025, PASS on disparate impact, governance and risk assessment; among disclosed groups impact ratio runs 0.914–1.000, with N/A for three small groups. The report states "Testing conducted by: Harver"SourceNone. As of 4 August 2026 we have no formal bias audit and no adverse impact ratio report
What is published about validityA peer-reviewed ACM FAccT 2021 paper on fairness mechanics, co-authored with four company employees and funded by the vendor; its Out of Scope section states the auditors did not investigate whether the games measure human capabilities or map to job performance. The product page carries the vendor's own figures: 98% completion rate, 59% reduction in time to hireSourcePsychometrics published on this site: CFI 0.96, Cronbach's α 0.69–0.77, test–retest above 0.83, R² 0.46 and AUC 0.77 on a validation sample of 3,000+, ML accuracy 72–89%. No peer-reviewed publication of our own
Languages"Translated into over 20 languages", with built-in adaptations for colour blindness and dyslexia, per the FAccT 2021 paper (2020–2021 data). Harver's acquisition press release of 11 August 2022 uses the phrase "compliant in 100+ countries, 30 languages"; the release gives no product-level breakdown, so we report it as a release-level claim rather than a pymetrics figure. We found no more recent public statement of the language list as of 4 August 2026SourceInterface and report in Russian and English; the client-facing profile handbook exists only in Russian
MENA presence and ArabicWe did not find a MENA breakdown in the vendor's public materials as of 4 August 2026; the acquisition press release (11 August 2022) uses the phrase "compliant in 100+ countries, 30 languages", without listing the countries or languages and without a product-level breakdown. We were unable to confirm either the presence or the absence of an Arabic localisationSourceCompany based in the UAE, serving clients in the region; no Arabic localisation of the product at present
Handling of demographic dataPer the FAccT 2021 audit, demographic data were used for evaluation only and not as model features; the vendor reported to the auditors that over 75% of players complete the optional demographic survey — a vendor-reported figure the auditors classed as not auditableSourceThe client receives anonymised codes and maps them itself; name, gender and age never reach NeuroFrame, so the model never receives protected attributes as inputs — an architectural constraint, not a guarantee that bias is absent. The cost of that scheme is the absence of ATS connectors
OwnerHarver. The acquisition was announced on 11 August 2022 and the financial terms were not disclosed; the pymetrics name is still used for the product lineSourceNeuroFrame; no change of ownership

pymetrics is not a standalone company: it is a product line inside Harver. The forms used on this page are "pymetrics (Harver)", "Harver" and "the pymetrics product line"; we do not write "the company pymetrics" in the present tense, and we do not treat "Pymetrics Inc." as a current vendor. pymetrics and Harver are trademarks of their respective owners, named here nominatively — to identify the product itself — without logos or brand styling.

What pymetrics is today

As of 4 August 2026, pymetrics is not an independent company. Harver announced the acquisition on 11 August 2022; the financial terms were not disclosed (harver.com/press/harver-acquires-pymetrics/). The name is still in use: Harver's own materials describe the product line as "pymetrics for Game-Based Assessments", and the product page at harver.com/gamified-assessments/ still carries the pymetrics brand.

The same acquisition release states that Harver had processed more than 100 million candidates and served more than 1,300 clients, and names AB InBev, Booking.com, Peloton, Valvoline and McDonald's. These are figures for Harver as a whole. Nothing in the release attributes them to the game-based assessment specifically, and we have not found a separate published figure for pymetrics volume.

Product work continued after the acquisition. A Harver press release dated 1 November 2022 describes customisable capabilities and values inside pymetrics assessments (harver.com/press/pymetrics-customizable-capabilities/). One naming detail is worth recording as a fact rather than a trend: the 2025 audit of this line, which we read in full, is titled "Harver's Soft Skills Platform", while the 2023 audit of the same line is recorded as "pymetrics inc. (Harver) Soft Skills Platform" in the ACLU's crowd-sourced register of Local Law 144 audits, which dates it 29 June 2023 and names BNP Paribas as the deploying employer. The employer-hosted copy of that 2023 report we have a link to (Paramount) returns HTTP 403 as of 4 August 2026, so its title here comes from the register rather than from the document. We draw no conclusion from that about the future of the brand — the product page still uses it.

On 21 April 2026 Harver announced AI PREVAIL, a non-game assessment of AI readiness with four dimensions (Business Wire, 21 April 2026). The release states "Trusted by more than 1000 enterprise organizations globally" and cites a 25% reduction in 90-day attrition, up to 50% faster time-to-hire and more than 60% reduction in overall attrition; the release attributes all of those to the Harver platform as a whole, and the word pymetrics does not appear in it. We draw no conclusion from that about the vendor's product priorities.

For context on where candidates meet the product: university careers services publish walkthroughs of the games — for example, a page on the Oxford University Careers Service site dated 24 November 2021. That page was written by an employee of JobTestPrep, a commercial test-preparation company, so it is a prep-vendor guest article hosted on a university site rather than institutional expertise.

How the assessment is built: format, number of tasks, duration

The vendor's own product page states, verbatim, "12+ interactive, gamified experiences" and "25 Minutes to complete" (harver.com/gamified-assessments/, checked 4 August 2026). The exact current number of games is not specified there — the page says "12+".

The peer-reviewed ACM FAccT 2021 paper, written with privileged access to the system, describes "a core set of twelve games that are derived from peer-reviewed psychological studies". The same paper states that the games run in a browser and a mobile app, are "translated into over 20 languages", include built-in adaptations for players with colour blindness and/or dyslexia, and had a gameplay base of over 2 million users. Those figures describe the product as of summer 2020.

That paper needs one caveat every time it is cited. Four of its eight authors were pymetrics employees, including then-CEO Frida Polli, and the audit was funded by the company through a $104,465 grant to Northeastern University (MIT Technology Review, 11 February 2021). We read it as a joint publication by auditors and the vendor rather than as independent research, and we state the reason above.

The twelve games are named publicly: Balloon, Tower, Money Exchange #1 and #2, Keypress, Hard or Easy, Digits Memory, Stop, Arrows, Lengths, Cards and Faces (Oxford University Careers Service, 24 November 2021 — the prep-vendor article noted above). The often-quoted figure of one to three minutes per game comes from strategycase.com, a commercial consulting-prep firm; we could not find a per-game breakdown in Harver's own documentation. The vendor's page lists nine top-level attributes: Effort, Risk Tolerance, Decision making, Attention, Focus, Learning, Fairness, Generosity and Emotion.

The scoring architecture, as described in the FAccT paper: a model is trained for each client. The "in group" is typically 50–100 high-performing employees of that client; the search for a fair feature set is run against impact ratio at the 70th percentile; final models are tested at the 50th and 70th; adverse-impact checking uses a separate bias group typically exceeding 10,000 users. The audited code used a proprietary SVM implementation over 64 features, and produced tiers — Not Recommended, Recommended, Highly Recommended.

What each side publishes about validity and reliability

What follows is what we were able to find published, taken by subject: fairness mechanics first, predictive validity second. On fairness mechanics the published material we found is the FAccT 2021 paper written with source-code access and the BABL AI bias audits covered in the next section. On the product page the vendor publishes its own performance claims — a 98% completion rate, a 59% reduction in time to hire, and the line "Validated for fairness across gender, ethnicity, and socioeconomic status". Those are the company's own figures; we found no independent confirmation of them in open sources.

The FAccT audit is explicit about what it did not examine. In the Out of Scope section, verbatim: the auditors "did not investigate the ability of pymetrics' games to measure human capabilities, whether those capabilities map to job performance, or whether other assessment methods would be superior". That is a boundary of the audit's mandate, not a finding that such validity is absent.

Our own search result, stated as such: as of 4 August 2026 we did not find, in open access, peer-reviewed publications on the predictive validity of pymetrics against job performance, nor a publicly posted technical manual. That is a statement about what our search returned, not about what exists. The FAccT paper itself refers to internal studies supplied to the auditors and to ongoing longitudinal analysis, neither of which is public.

One frequently misused data point deserves its methodology. Raghavan, Barocas, Kleinberg and Levy (2019) reviewed 18 algorithmic hiring vendors by "exhaustively searching their websites, downloading whitepapers they provide, and watching webinars", and recorded pymetrics' validation disclosure in that table as a "small case study". That is a snapshot of what was public on the vendor's website in 2019, three years before the Harver acquisition. It describes the state of public disclosure at that time and cannot be used to characterise the product's evidence base today. The same authors raise a separate methodological point, addressed to the industry rather than to any one vendor: where a vendor advertises "millions of data points" per candidate, they argue it may be impossible to give a substantive rationale for every included feature, while the APA Principles require that a clear rationale be established linking scores to the criterion constructs of interest. Whether that description fits the product discussed here we do not know and do not assert.

For symmetry, here is what we publish about ourselves. Construct validity: CFI 0.96. Internal consistency: Cronbach's α 0.69–0.77. Test–retest reliability above 0.83. Link to work outcomes: R² 0.46 and AUC 0.77 on a validation sample of 3,000+ against real KPIs and manager ratings; ML model accuracy 72–89%. We deliberately do not place our 0.46 next to meta-analytic coefficients from other instruments: the samples and criteria differ, and the comparison would not survive scrutiny. And we have no peer-reviewed publication of our own validation studies at all — see the section on when to choose them.

Bias audits: what was published, by whom, and what it covers

First, a point that is routinely stated backwards. Under NYC Local Law 144 the duty to publish the summary of bias audit results falls on the employer or employment agency, not on the assessment vendor: DCWP rule §5-303 (Published Results) is addressed to "an employer or employment agency"; §5-301 sets out how the audit itself must be calculated. A vendor that publishes an audit on its own site is publishing something the rule text addresses to employers and employment agencies; whether and how the rule reaches a particular vendor is for that vendor and its counsel to state.

Harver publishes one. BABL AI Inc. (Iowa City, Iowa) issued a Summary of Bias Audit Results, V1.0, dated 17 July 2025, for a system named "Harver's Soft Skills Platform". The data period is January–December 2024; the report states "Testing conducted by: Harver" and "Date of most recent testing: Jun 2025"; the selection rate is computed at the 50th percentile threshold. The disclosed figures, as N / selection rate / impact ratio: Male 164,014 / 0.545 / 1.000; Female 119,242 / 0.540 / 0.992; Black or African American 16,908 / 0.590 / 1.000; White 63,086 / 0.571 / 0.967; Hispanic or Latino 15,030 / 0.542 / 0.918; Asian 58,958 / 0.539 / 0.914. Intersectional groups are also disclosed, with the lowest at Hispanic/Latino Male, 7,891 / 0.533 / 0.917. Among the groups for which a ratio was calculated, all are above the four-fifths threshold. The opinion is PASS on disparate impact, governance and risk assessment.

The report states its own limits, and they matter as much as the numbers. It does not certify the model as "bias-free"; it did not provide sufficient testing of disparate impact on any protected class beyond race/ethnicity and gender; and it does not determine whether the model is, in fact, an AEDT within the meaning of Local Law 144. Three small groups — Native Hawaiian or Pacific Islander (N=378), Native American or Alaskan Native (N=398) and Two or more races (N=2,440) — carry an impact ratio of N/A. And the report discloses that 331,177 candidates with at least one unknown demographic attribute were excluded from the calculations, alongside 191,455 with unknown gender and 317,513 with unknown race/ethnicity. Disclosure of that count is what the DCWP rules ask for: §5-301(b)(5) has the audit "indicate the number of individuals the AEDT assessed that are not included in the required calculations because they fall within an unknown category", and §5-303(a)(1) puts that number into the summary the employer publishes. How such records are to be treated inside the calculation itself the rules do not prescribe.

An earlier audit cycle exists. The ACLU's crowd-sourced register of Local Law 144 audits records a BABL AI audit of the "pymetrics inc. (Harver) Soft Skills Platform" dated 29 June 2023, published by an employer — BNP Paribas — rather than by the vendor, as the law contemplates. We have not read that document: as of 4 August 2026 both the employer-hosted copy we have a link to (Paramount) and the archive capture listed in the register return HTTP 403, so everything we say about it comes from the register entry. That publication pattern is also the reason audits of game-based assessments are hard for a candidate to locate: they often live on the employer's careers page rather than the vendor's.

Industry-wide context for reading any published audit comes from "Null Compliance" (arXiv:2406.01399): 18 of 391 employers published audit reports, and of 386 individual impact ratio values appearing in 13 unique reports, 377 (96%) were above 0.8 — a distribution the authors attribute in part to publication bias. Neither pymetrics nor Harver is mentioned in that study, and its findings cannot be transferred to them.

The regulatory map as of August 2026

First, the date, which is often quoted from the Act's original text. The EU AI Act does classify AI systems used to evaluate candidates in recruitment as high-risk — Annex III, point 4(a). But the obligations for Annex III high-risk systems no longer start on 2 August 2026: Regulation (EU) 2026/1744 of 8 July 2026, published in the Official Journal on 24 July 2026, moved the date of application of Chapter III, Sections 1–3 to 2 December 2027 for Annex III systems and to 2 August 2028 for Annex I systems. The dates are fixed, with no conditional triggers.

The deferral is partial, not general. The prohibitions in Chapter II were not softened: since 2 February 2025 the Act has prohibited AI systems used to infer the emotions of a natural person in the workplace, outside medical and safety grounds, and the Commission has published guidance on that prohibition. On our reading of that guidance the prohibition is tied to inferences drawn from biometric data, and assessment that runs on behavioural telemetry — clicks, timings, choices — is a different case. That is how we read the text, not a legal conclusion about any product, ours or anyone else's. Most Article 50 transparency obligations remain at 2 August 2026. We make no assessment of any vendor's position under any of these provisions; that is for the vendor and its counsel to state.

GDPR is the older and, for European hiring, the more immediate constraint. Article 22 gives a person the right not to be subject to a decision based solely on automated processing that significantly affects them, and requires at minimum a right to human intervention, to express a point of view and to contest the decision. Two CJEU rulings sharpen this: SCHUFA (C-634/21, 7 December 2023), under which Article 22 duties can reach the score provider and not only the employer where the recipient strongly draws on the score; and Dun & Bradstreet Austria (C-203/22, 27 February 2025), under which trade secrecy cannot as a rule block the explanation a person needs in order to contest a decision.

In the United States the four-fifths rule of UGESP (29 CFR §1607.4(D)) is a rule of evidence, not a safe harbour: a ratio above 0.8 "will generally not be regarded" as evidence of adverse impact, and the same passage adds that smaller differences may still constitute adverse impact where they are significant in both statistical and practical terms. Federal EEOC guidance on AI in hiring was withdrawn in 2025 and enforcement priority has shifted towards disparate treatment, but Title VII and UGESP remain in force, and private and state-level exposure is unchanged.

For the Gulf, where our own clients mostly sit: there is no dedicated AI-in-hiring statute. The binding instrument at federal UAE level is Article 18 of the PDPL (Federal Decree-Law No. 45 of 2021), giving a data subject the right to object to decisions issued by automated processing, including profiling. DIFC Regulation 10, adopted 1 September 2023, goes further than anything else in the region: it requires disclosure that a system is in use, documentation of bias-detection and human-intervention mechanisms, and — for high-risk processing — certification and a designated Autonomous Systems Officer. The UAE AI Charter of 2024 is advisory.

What the provisions cited above have in common is that we found no prohibition of game-based assessment in their texts, and that requirements on explainability, documented bias checking and human intervention do appear in them. That is a description of the texts, not a legal opinion: how they apply to a given deployment is for your counsel to determine.

Two ways to set a norm: your own high performers, or an external reference

Every assessment answers the question "compared with what?", and there are two families of answer. Either the standard is derived from people who already work in the role at this specific employer, or the standard is built outside the client and the candidate is compared against it. This is the substantive difference between the two products, and it is a trade-off in both directions rather than a ranking.

pymetrics sits in the first family. As described in the FAccT 2021 paper, a model is trained per client from an in group of typically 50–100 high-performing employees, contrasted with an out group, and screened for adverse impact on a separate bias group typically exceeding 10,000 users. The audit found that demographic data were used for evaluation only and not as model features, and — a finding that favours the vendor and is worth quoting — that when the auditors submitted a deliberately poisoned in group, "we were unable to produce a biased model that was not flagged as such by the code".

The structural criticism of incumbent-based training is long-standing and is directed at the approach as a class, not at any one product: Raghavan et al. (2019) raise it in their review of 18 vendors, and MIT Technology Review (21 July 2021) reports disability advocates' concerns about standards derived from current successful employees. Whether and how any particular implementation addresses it is a question for the vendor.

NeuroFrame sits in the second family. Our reference is built outside the client: 287 professions across 8 corporate lifecycle stages, giving 2,296 benchmark cells, derived from the Russian Ministry of Labour occupational register and professional standards together with the Adizes lifecycle model. The comparison sample is 14,850 people in real jobs across 21 industries, 21 functional areas and 8 grades, drawn from 500 companies. Two consequences follow mechanically: the method works from day one, without requiring the client to have dozens of current employees in the role, and it does not structurally encode the existing composition of the workforce.

The reference is stage-dependent, and that is its point. On the 2,296 records, for 99% of the 287 professions the requirement ranges at the Infancy and Bureaucracy stages overlap by less than half, with a median range overlap of 0.40. Of the 153 non-intersections between the extreme stages, 152 fall on risk appetite. Across the lifecycle the centre of the range moves by 18.8 percentile points for risk appetite, 15.4 for openness, 11.5 for conscientiousness and 9.0 for cognition; result focus barely moves at all.

The honest counterweight: an external reference does not know your company. It cannot tell you what specifically works in your culture, in your product, under your management — and that is precisely what a model trained on your own high performers is built to capture. If you have a large, stable, well-measured cohort in a role, that is a real argument in its favour.

What NeuroFrame does differently

NeuroFrame is a single continuous simulation lasting 30–60 minutes, not a battery. It yields eight parameters — three in the cognition domain (Mental Efficiency, Learning Agility, Progress Monitoring) and five in personality (Openness to New, Risk Appetite, Result Focus, Agreeableness, Conscientiousness). It issues no hire/no-hire verdict: the output is a ranking and a fit-to-role reading, and the decision stays with the employer. That is a stated policy constraint, not a limitation we discovered later.

Every benchmark cell carries a human-readable rationale. All 2,296 of them: 1.07 million characters in total, 464 characters per cell on average, and 32% of them name the rule that fired. This is what the 2026 regulatory environment — Article 22 GDPR, the CJEU explanation rulings, DIFC Regulation 10, the coming Annex III regime — keeps asking of systems that support personnel decisions. In 303 of the 2,296 cells (13%) the benchmark permits a high risk appetite alongside a low floor on cognition or on progress monitoring: the method can say "fits on every parameter, and this particular combination warrants attention". We surface that as a separate line next to the parameters rather than folding it into one overall score: in our experience, collapsing to a single number is where such a combination disappears from the report.

Anonymity is built into how data reaches us. The client receives codes and maps them to people itself; NeuroFrame never receives a name, a gender or an age. The model therefore never receives protected attributes and cannot train on them directly. That is an architectural constraint, not a guarantee that bias is absent: indirect correlates remain possible, which is exactly why we state plainly that we have no adverse impact ratio report yet. The cost is direct and we name it in the next section: this scheme is why we have no ATS connectors.

On what the game format does and does not buy, we follow the literature rather than our own marketing. A systematic review of 34 studies (Ramos-Villagrasa, Fernández-del-Río and Castro, Frontiers in Psychology, 2022) concludes that game-based formats do not offer sufficient advantages to be recommended over conventional methods, except in improved candidate reactions — so the defensible claim is comparable validity with substantially better candidate experience, and nothing stronger. Vigilance decrement — the decline of sustained attention with time on task — is one of the most reproducible effects in cognitive psychology at the group level; a long session therefore lets you observe behaviour both fresh and fatigued, which is not the same as scoring an individual's accumulated fatigue, and we do not claim the latter.

One construct argument is worth making precisely. Dynamic simulations with delayed consequences measure complex problem solving, a construct that overlaps general cognitive ability by roughly 18% of variance (r ≈ .43 across 47 studies — Stadler, Becker, Gödker, Leutner and Greiff, Intelligence, 2015) — that is, it contains something classic tests do not capture. It is also worth keeping the current selection-method coefficients in view: on the Sackett, Zhang, Berry and Lievens (2022) recalculation, structured interview stands at .42, job knowledge tests at .40, work samples at .33, cognitive ability tests at .31, assessment centres at .29, situational judgement tests at .26 and conscientiousness questionnaires at .19. We do not think any assessment, ours included, replaces a structured interview. The ranges are a reference point, not a verdict: a gap on one or two parameters is something to explore at interview, not grounds for rejection.

One long task versus twelve short ones: what actually follows

A discrete, named task can be found, discussed and rehearsed; a continuous scenario is harder to decompose into moves. This is a property of the format, not a claim about the quality of any product — and it cuts in both directions, as the rest of this section shows.

The observable evidence is the preparation industry that forms around named tasks. The walkthrough of the pymetrics games hosted by the Oxford University Careers Service was written by an employee of JobTestPrep (24 November 2021); strategycase.com sells preparation for consulting assessments. That is the existence of a preparation market, not a claim that preparation changes anyone's score — we found no data either way.

The battery format buys breadth, and here the advantage is theirs. Twelve games built on classical cognitive paradigms cover more distinct constructs than a single scenario, and Harver additionally sells separate cognitive tests — perceptual speed, verbal, spatial, logical and mathematical reasoning — as add-ons outside the twelve games (per GraduatesFirst, a commercial prep source; we did not find a vendor page enumerating them). NeuroFrame does not measure verbal or numerical ability at all. If you need those, you need something else, whether alongside us or instead of us.

Time is the other honest trade. The vendor states 25 minutes to complete; NeuroFrame takes 30 to 60. At the scale of a graduate intake of tens of thousands of applicants, a 20-minute difference in candidate time is a real operational number, and it is not in our favour. Our answer is not that our session is short — it is that a single continuous session is where delayed consequences, changing conditions and behaviour under accumulating time-on-task become observable at all. Whether that is worth the extra minutes depends on what you are hiring for.

How to read this page, and how to check it

This comparison is published by NeuroFrame — that is, by an interested party. That is why every fact about the other product carries a link you can open — a primary source wherever one is available, and, in the two places where it is not, the secondary source we could open, labelled as such — and why we have tried to avoid grading the other product. Adjectives like "weak", "shallow" or "outdated" are deliberately kept off this page; if one has slipped through, tell us and we will remove it.

The rules we worked to: vendor claims come from the vendor's own pages and press releases and are labelled as the vendor's claims; audit figures come from the audit PDF and are labelled with who performed the testing; peer-reviewed material is cited with its conflicts stated; and commercial test-preparation sites are marked as such every time they appear, because they are not primary sources about the product. Everything was checked on 4 August 2026, and vendor pages change — the date tells you what we saw.

What we deliberately did not write. We did not write that the vendor lacks validity evidence: not finding something is not the same as it not existing, and the FAccT paper itself refers to internal studies and ongoing longitudinal analysis. We did not write anything about the vendor's compliance posture under any regulation. We did not write anything about the procurement decisions of the vendor's named clients — the widely circulated claim about BCG, for instance, traces to a commercial prep site describing an office-dependent situation and a direction of travel, which is not a statement by BCG or by Harver. And we did not mix Harver's platform-wide numbers with pymetrics' numbers, because we found no source that connects them.

If you find an error on this page, it is an error we want to fix rather than defend. Our aim is that every number here is either sourced to a URL you can open or, for our own figures, traceable to the single internal source of truth the whole site is generated from — and where a number falls short of that, we would rather hear about it than leave it standing.

When to choose pymetrics (Harver) rather than us

If you are a US employer, and specifically a New York City one, choose them. Under Local Law 144 the obligation to publish a bias audit is yours, not the vendor's — and a vendor that already has a published BABL AI audit dated 17 July 2025, with selection rates and impact ratios by gender, race and intersection, hands you a starting point. We have no formal bias audit and no adverse impact ratio report as of 4 August 2026. That is the single clearest gap between us, and it is not one you should discount.

If you need peer-reviewed evidence, choose them. There is a peer-reviewed ACM FAccT paper about their fairness mechanics, with all the co-authorship caveats we set out above. We have no peer-reviewed publication of our own validation studies at all. Our psychometrics are published on this site and traceable to an internal source of truth, and that is a weaker form of evidence than a journal.

If your hiring is multilingual or global, choose them. The FAccT paper documents translation into over 20 languages plus built-in adaptations for colour blindness and dyslexia. Our client-facing profile handbook exists only in Russian, our accommodations policy for candidates with disabilities is not yet formalised, and our compliance perimeter is Russian: 152-FZ is closed architecturally, but we do not yet have a GDPR, EU AI Act or NYC LL 144 package.

If you screen at high volume through an ATS, choose them. We have no ATS connectors — a direct consequence of the anonymised-code scheme we described, and a real operational cost, not a philosophical position. Add the time difference: 25 declared minutes against our 30 to 60.

If you need breadth of constructs, choose them. Twelve games and — per the commercial prep source GraduatesFirst, since we found no vendor page listing them — separate cognitive tests covering verbal, numerical, logical and spatial reasoning; we run one task and do not measure verbal or numerical ability at all.

If you are hiring senior executives, look carefully at us before committing. 83% of our reference library sits at specialist level, and top management is covered by nine professions. And if you already have a large, stable cohort of proven high performers in the role you are filling, a model trained on those people has an argument we structurally cannot make.

Two more numbers of ours you should weigh rather than take on trust. Our internal consistency is α 0.69–0.77, below the conventional 0.80 threshold. Our normative base is 14,850 people — smaller than instruments with forty-year histories. We publish both because a page that only lists its own strengths is an advertisement, and you should read it as one.

Questions

Does pymetrics still exist, or is it Harver now?
Both. There is no independent pymetrics company: Harver announced the acquisition on 11 August 2022, with financial terms undisclosed. But the name is alive and used by the vendor — the product line is called "pymetrics for Game-Based Assessments" and the product page at harver.com/gamified-assessments/ carries the brand. One detail for accuracy: the 2025 audit is titled "Harver's Soft Skills Platform", while the 2023 audit is recorded as "pymetrics inc. (Harver) Soft Skills Platform" in the ACLU's crowd-sourced register of Local Law 144 audits — the 2023 document itself does not open at any link known to us as of 4 August 2026 (HTTP 403), so its title here comes from that register. That is a fact about document titles, and we draw no conclusion from it about the brand's future.
How long does it take a candidate — and why are you longer?
The vendor states 25 minutes to complete (harver.com/gamified-assessments/, checked 4 August 2026). NeuroFrame takes 30–60 minutes. At high volume that difference is not in our favour, and we say so. The point of a long session is different: delayed consequences, changing conditions and behaviour under accumulating time-on-task are observable only inside one continuous session. If your priority is minimum candidate time in volume screening, that is an argument for them.
They have a published bias audit and you don't. How much does that matter?
It matters, and we won't soften it. Harver has a published BABL AI report dated 17 July 2025 with a PASS opinion and disclosed impact ratios; we have no formal audit. Two clarifications on the substance, not in mitigation. First, under NYC Local Law 144 the duty to publish falls on the employer, not the vendor, so a vendor-published audit is voluntary — but it does hand you a starting point. Second, the BABL AI report itself states that it does not certify the model as "bias-free", did not test protected classes beyond race/ethnicity and gender, and does not determine whether the system is an AEDT under the law. An audit is a documented check, not a certificate that bias is absent.
What does "a model trained on my own employees" mean, and what if I don't have 50 people in the role?
As described in the FAccT 2021 paper, a pymetrics model is built per client: an in group of typically 50–100 high performers, an out group for contrast, and adverse-impact screening on a separate sample of 10,000+. If you have such a cohort, and it is stable and well measured, that is a strong argument — the model captures what works specifically at your company. If you don't — a new role, a startup, a new market — that route is closed. Our reference is built externally, from an occupational standards register and the Adizes model, so it works from day one and does not require incumbents in the role. The cost is symmetrical: an external benchmark does not know your culture.
Which of you has demonstrated predictive validity?
Neither answer here is comfortable. On pymetrics: as of 4 August 2026 we did not find, in open access, peer-reviewed publications on predictive validity against job performance, nor a publicly posted technical manual — that is a statement about our search, not about what exists, and the FAccT paper itself refers to internal studies given to the auditors and to ongoing longitudinal analysis. On NeuroFrame: we publish R² 0.46 and AUC 0.77 on a validation sample of 3,000+ against real KPIs and manager ratings, but we have no peer-reviewed publication either. And we do not place our 0.46 next to other instruments' meta-analytic coefficients: the samples and criteria differ.
We hire in the UAE. What is actually regulated here?
There is no dedicated AI-in-hiring law in the UAE. The binding federal provision is Article 18 of the PDPL (Federal Decree-Law No. 45 of 2021): a candidate has the right to object to a decision issued by automated processing, including profiling. In the DIFC, Regulation 10 applies (from 1 September 2023) — it requires disclosing that a system is in use, documenting bias-detection and human-intervention mechanisms, and, for high-risk processing, certification plus a designated Autonomous Systems Officer. The UAE AI Charter of 2024 is advisory. We found no prohibition of game-based assessment in the texts of these provisions, and requirements on explainability and human intervention do appear in them. That is a description of the texts, not a legal opinion: how they apply to your deployment is for your counsel to determine.
Is it true that the EU AI Act already requires all this from August 2026?
No — the date is often quoted from the Act's original text, before the deferral. Systems used to evaluate candidates are indeed classified as high-risk (Annex III, point 4(a)), but their obligations have been deferred: Regulation (EU) 2026/1744 of 8 July 2026, published in the Official Journal on 24 July 2026, set the date of application of Chapter III, Sections 1–3 to 2 December 2027 for Annex III systems. The deferral is partial, though: the prohibition on inferring a person's emotions in the workplace (Art. 5(1)(f)) has applied since 2 February 2025 and is limited, per Commission guidance, to inferences from biometric data, while Article 50 transparency obligations remain at 2 August 2026.

Sources

Every link was opened and checked on the date shown above.

  1. Harver Acquires pymetrics — official press releaseHarver
  2. Harver Acquires Pymetrics: One Team, One MissionHarver
  3. pymetrics Gamified Assessments — Game-Based Behavioral AssessmentsHarver
  4. Custom Options Added to pymetrics Assessments (press release, 01.11.2022)Harver
  5. Summary of Bias Audit Results — Audit of Harver's Soft Skills Platform for New York City's Local Law 144, V1.0, 17.07.2025BABL AI Inc.
  6. Summary of Bias Audit Results, 29.06.2023 — копия, размещённая работодателем; на 04.08.2026 ссылка отдаёт HTTP 403, документ нами не прочитанParamount / BABL AI
  7. Tracking Automated Employment Decision Tool Bias Audits — краудсорсинговый реестр аудитов NYC LL 144 (строка: BNP Paribas / «pymetrics inc. (Harver) Soft Skills Platform» / BABL AI / 29.06.2023)American Civil Liberties Union
  8. Wilson, Ghosh, Jiang, Mislove, Baker, Szary, Trindel, Polli — Building and Auditing Fair Algorithms: A Case Study in Candidate Screening (ACM FAccT 2021)ACM FAccT 2021
  9. Raghavan, Barocas, Kleinberg, Levy — Mitigating Bias in Algorithmic Hiring: Evaluating Claims and Practices (2019)arXiv
  10. Auditors are testing hiring algorithms for bias, but there's no easy fix (11.02.2021)MIT Technology Review
  11. Disability rights advocates are worried about discrimination in AI hiring tools (21.07.2021)MIT Technology Review
  12. Null Compliance: NYC Local Law 144 and the Challenges of Algorithm AccountabilityarXiv
  13. Notice of Adoption of Final Rule — Automated Employment Decision Tools, 6 RCNY §§ 5-300–5-304NYC Department of Consumer and Worker Protection
  14. Regulation (EU) 2026/1744 of 8 July 2026 (Digital Omnibus on AI), OJ 24.07.2026EUR-Lex / Official Journal of the European Union
  15. EU AI Act, Annex III (high-risk use cases), point 4(a)artificialintelligenceact.eu
  16. Red Lines Under the EU AI Act: the prohibition of emotion recognition in the workplace and educationFuture of Privacy Forum
  17. Art. 22 GDPR — Automated individual decision-making, including profilinggdpr-info.eu
  18. Key takeaways from the CJEU's automated decision-making rulings (SCHUFA, C-634/21)IAPP
  19. Court of Justice press release No 22/25 — C-203/22 Dun & Bradstreet Austria (27.02.2025)Court of Justice of the European Union
  20. 29 CFR § 1607.4 — Uniform Guidelines on Employee Selection Procedures, four-fifths ruleCornell Legal Information Institute
  21. The federal government quietly removed its AI hiring guidanceNational Law Review
  22. AI in the UAE: understanding the regulatory landscape and key authoritiesLatham & Watkins
  23. AI regulation in the DIFC: personal data processed through autonomous and semi-autonomous systems (Regulation 10)Mayer Brown
  24. Harver Launches AI PREVAIL — AI Readiness Assessment (Business Wire, 21.04.2026)Business Wire
  25. The Pymetrics Games — Overview and Practice Guidelines (24.11.2021, guest article by a JobTestPrep author)Oxford University Careers Service
  26. BCG Pymetrics Test 2026: Is It Still Used? (commercial preparation site)StrategyCase
  27. Harver/Pymetrics Assessments & Digital Interviews (commercial preparation site)GraduatesFirst
  28. Sackett, Zhang, Berry & Lievens — Revisiting meta-analytic estimates of validity in personnel selection (2022), Journal of Applied Psychology 107(11), 2040–2068American Psychological Association
  29. Ramos-Villagrasa, Fernández-del-Río & Castro — Game-related assessments for personnel selection: A systematic review (2022), Frontiers in Psychology 13, 952002 (обзор 34 работ)Frontiers in Psychology
  30. Stadler, Becker, Gödker, Leutner & Greiff — Complex problem solving and intelligence: A meta-analysis (2015), Intelligence 53, 92–101 (47 исследований, средний эффект .433)Elsevier / Intelligence