Comparison
NeuroFrame and Arctic Shores
This comparison is published by NeuroFrame — that is, by an interested party. So every fact about the other product carries a link to a primary source, and we assign no ratings at all. Arctic Shores is a British assessment vendor whose product is a battery of short interactive tasks; NeuroFrame is a single continuous simulation benchmarked against an externally built reference library. Below is what each side publishes, checked on 4 August 2026, and an honest section on where Arctic Shores is the better fit.
Updated:
Side by side
| Attribute | Arctic Shores | NeuroFrame |
|---|---|---|
| Owner | Arctic Shores Limited, private UK company, number 08589048, incorporated 28 June 2013, registered in Manchester; status Active per Companies House on 4 August 2026. We found no record of an acquisition.Source | NeuroFrame — an independent product, not part of a larger vendor group. |
| Instrument class | A battery of short interactive tasks. The site promotes it as a task-based assessment; the glossary section of the same domain describes it as "our game-based assessment" (checked 4 August 2026).Source | One continuous strategic simulation with delayed consequences. No separate named mini-games. |
| Number of tasks | Not fixed. The vendor brochure says "a series of up to eight engaging tasks" (link in Sources); the candidate guidance page linked here says "Up to nine tasks will encourage and capture your unique behaviour" (checked 4 August 2026). Client candidate material states that composition varies with the role (Teach First, checked 4 August 2026).Source | One task — a single simulation. No subtest battery; verbal and numerical ability are not covered. |
| Duration | Vendor: "Know true potential, in 45 minutes" (How it works) and "All in under 45 mins" (brochure) — both links in Sources. Client candidate material states a wider window: "anywhere between 15-50 minutes to complete the assessment" (Teach First, checked 4 August 2026; link in this row).Source | 30–60 minutes in a single uninterrupted session. |
| Published psychometrics | On the four public pages we checked on 4 August 2026 (homepage, How it works, The Science Behind Our Assessments, Compliance) no numerical reliability or validity coefficients and no sample sizes appear. Technical manuals by the vendor exist and are cited in the academic literature; we did not find them in open access.Source | Cronbach's α 0.69–0.77; test–retest above 0.83; CFI 0.96; R² 0.46 and AUC 0.77 on a validation sample above 3,000. No peer-reviewed publication of our own validation studies. |
| Published bias audit | The Compliance page (checked 4 August 2026) names GDPR, the Data Protection Act 2018 and ISO 27001 certification. Employer client FDM Group publishes an NYC applicant bias audit page listing Arctic Shores among its AEDT suppliers; we could not verify the contents of the linked report.Source | No formal bias audit and no adverse impact ratio report. The platform receives no name, gender or age by design, so the model cannot train on protected characteristics — a design property, not a substitute for an audit. |
| Languages | We did not find a published list of interface languages in open sources as of 4 August 2026, and found no separate confirmation of an Arabic localisation. The vendor's Series B release of 16 January 2023 states "more than three million candidates worldwide" across "over 40 countries" (checked 4 August 2026).Source | Candidate experience and report in Russian and English. The client-facing profile handbook exists only in Russian. |
| Presence in MENA | In the current case-study list (16 entries, checked 4 August 2026) we found no Middle Eastern clients; the geography shown there is the UK and Europe, including Greece. The KPMG UAE case page referenced by earlier materials returns HTTP 404 as of 4 August 2026. The reason is not known to us.Source | Present in the UAE and Russia. The Russian data perimeter (152-FZ) is closed architecturally; a GDPR, EU AI Act and NYC LL 144 package is not yet in place. |
| Norming model | Results are compared against norm groups; decisions are taken on a composite fit score rather than on individual traits. The undated University Pack describing Skyrise City states: "Our clients never make decisions based on individual traits – this is why we created the fit score" and "Employers tend to deselect below the 50th percentile on our fit scores".Source | The requirement benchmark is built externally — from the Russian Ministry of Labour occupational register and the Adizes lifecycle model: 287 professions × 8 stages = 2,296 cells, each with a written justification. Comparison sample: 14,850 people. The model is not trained on the client's own staff. |
Arctic Shores Limited (United Kingdom, company number 08589048). All names and trademarks belong to their respective owners and are used nominatively, to refer to the product itself, without logos or brand styling.
Who wrote this, and under what rules
This page is written by a competitor. That is the single most important thing to know before reading it, and it is why the page is built the way it is: every claim about Arctic Shores is either a quotation from the vendor's own material or an entry in a public register, with a link and a date. Where we have no source, we have no sentence.
Three rules constrain what follows. We use no evaluative adjectives about the other product — no "weak", no "shallow", no "insufficient". We do not convert a failed search into a fact: if we did not find something, we say we did not find it, on a given date, in the sources we checked, and not that it does not exist. And we do not speculate about intent — where the vendor's own materials use two different terms for the same product, we quote both and stop there.
One deliberate omission. The vendor's help centre at help.arcticshores.com did not respond to our automated requests on 4 August 2026, so we could not verify its wording ourselves. Nothing on this page rests on it, and nothing is quoted from it. Commercial terms are also absent: we found no published price list on either side's public pages as of 4 August 2026.
This page is for someone already in a procurement process with Arctic Shores who wants to understand the difference in kind, not the difference in volume. If you are looking for a verdict, this is the wrong page — the last two sections deliberately argue against us.
What Arctic Shores is
Arctic Shores is a British assessment vendor. Its product measures candidates through their behaviour in short interactive tasks rather than through answers to questions. The named clients on the vendor's current case-study list (16 entries, checked 4 August 2026) are predominantly UK and European organisations, and several of the cases describe graduate and volume-hiring programmes.
The legal entity is Arctic Shores Limited, company number 08589048, incorporated on 28 June 2013, registered office at Lowry House, 17 Marble Street, Manchester M2 3AW, status Active per Companies House as of 4 August 2026. We found no record of an acquisition: the company appears to remain independent.
On funding, the vendor announced a £5.75m Series B on 16 January 2023, led by Praetura Ventures and Calculus Capital with participation from existing shareholder Beringea; the same release stated "more than three million candidates worldwide" across "over 40 countries". An earlier $5.5m round from September 2019 led by Beringea, with Candy Ventures among the investors, is described in the vendor's own brochure — an undated document that, judging by its internal references, dates from roughly 2021–2022, so figures in it such as "Arctic Shores employs 60 people" should be read as of that period rather than as current.
Two changes are worth noting for anyone evaluating the vendor now. Leadership changed: Estelle McCartney was appointed CEO and co-founder Robert Newry moved to a role the announcement calls "Chief Explorer"; that page carries the date "Wednesday 22nd January" without a year. And the product itself moved twice — a construct model the vendor calls Skill-enablers™ launched on 17 September 2024, and a separate Learning Agility assessment was announced on 12 November, which the site's own sitemap dates to 2025.
How the assessment is built: tasks, count, duration
The short answer: neither the number of tasks nor the duration is fixed. Client candidate materials state that the composition varies by role — Teach First writes that "depending on the role you apply for, the first 1-2 tasks in the assessment may also be timed" and asks candidates to "set aside anywhere between 15-50 minutes to complete the assessment" (checked 4 August 2026). On the vendor's own public pages that we checked we found no corresponding statement. So a single number quoted elsewhere is best read as one configuration rather than as the product's specification.
On count, the vendor's own materials give two figures. The brochure describes "a series of up to eight engaging tasks". The candidate guidance page on the site says "Up to nine tasks will encourage and capture your unique behaviour" (checked 4 August 2026). The historical product, Skyrise City, is described in the vendor's University Pack as running across nine levels; the Birkbeck thesis that studied it reports participants completing eight. Read together, these are consistent with a configurable battery rather than with a contradiction.
On duration, the vendor's How it works page is headlined "Know true potential, in 45 minutes" and the brochure says "All in under 45 mins". Clients state a wider window in their own candidate materials: Teach First's page says "Please expect to set aside anywhere between 15-50 minutes to complete the assessment" and adds that "depending on the role you apply for, the first 1-2 tasks in the assessment may also be timed" (checked 4 August 2026). Both the client figure and the vendor figure can be true at once if the battery is configurable.
On output, the vendor describes a composite. How it works presents a "natural fit" score "from 1-100". The University Pack — undated, describing Skyrise City — is explicit about why: "Our clients never make decisions based on individual traits – this is why we created the fit score. If we gave a client all 30 measures to choose from there would be a huge amount of human error." The same document records employer practice at the time as "Employers tend to deselect below the 50th percentile on our fit scores".
One note on terminology, stated as fact and not interpreted. The site promotes the product as a task-based assessment, and on the same domain the glossary section describes it as "our game-based assessment" (checked 4 August 2026). We record both formulations and draw no conclusion about why they coexist.
What the vendor publishes about validity and reliability
The direct answer: on the four public pages we checked on 4 August 2026 — the homepage, How it works, The Science Behind Our Assessments, and Compliance — no numerical psychometric coefficients appear: no reliability estimates, no validity coefficients, no sample sizes. That is a statement about those four pages on that date and nothing more.
It is specifically not a claim that such data does not exist. Technical manuals by this vendor do exist and are cited in the academic literature: the reference list of the Birkbeck thesis discussed below contains entries for a Skyrise City Technical Manual (2019) and a Cosmic Cadet Technical Manual (2018). We did not find either document in open access on the vendor's site. For current psychometric data, the vendor is the right party to ask, and asking for it is a reasonable thing to do in procurement.
What the science page does carry, as of 4 August 2026, are claims in prose: that the team's job is "to keep it in line with BPS standards", that "our researchers conduct dozens of studies a year", and the assertion "no adverse impact, regardless of ethnicity or gender. Ever." We record that this statement is made. No supporting figures — sample sizes, impact ratios, or a study report — are published on that page itself.
Separately, the vendor publishes outcome and candidate-reaction figures on its marketing pages and in its case studies, without stating samples or significance tests. On the homepage: "85% of candidates enjoy taking our engaging assessment", "91% completion rate", "over 12,000 data points", "reduce turnover by 40% in just 6 months". In the brochure, an AXA case: hires "18% more likely to stay beyond a year", "41% fewer sick days", £2.5m over twelve months, 92% positive candidate feedback, 2,000–2,500 hires a year. On the case studies index: Entain turnover 40%→25%, Holcim UK women 3%→27%, Experis Academy 95% completion, Capita 12,000 applications. The Adecco entry carries a figure of 89% whose measured quantity is not stated in the index headline; we do not reproduce it without one. All of these are the vendor's own statements, attributed as such.
One more published item is relevant to anyone worried about AI-assisted cheating: on 3 October 2023 the vendor published an experiment in which, by its account, GPT-3.5, GPT-4, Google Bard and a home-built OCR bot were unable to complete its task-based assessment. The published account contains no numerical metrics.
What the academic literature says — and how to read it
The publicly available academic work on this instrument that we are aware of is a doctoral thesis: Liam Kevin Close, "The psychometric structure of a game-based assessment", PhD, Birkbeck, University of London, 2022. Before any of its findings, three qualifiers matter. It was carried out on Skyrise City — the previous generation of the product, not the current task-based assessment. It used data supplied by the vendor, with level images reproduced by permission, so it is not a third-party audit. And a thesis is not a peer-reviewed publication.
With those qualifiers stated, the method and the numbers. The author applied generalisability theory to samples of 398, 702 and 429 candidates. The abstract reports that aggregating to the level of individual dimensions produced "poor reliability estimates (0.36–0.66)", while the aggregated overall score "approached acceptable levels (0.63–0.86)". The recommendation in the text is that "reporting dimension-based scores is not recommended, as dimensions were found to contribute little variance to overall scores".
The thesis also preserves numbers from the vendor's non-public reports — we found them nowhere else in open access: a mean stratified internal consistency of 0.72 across traits (Arctic Shores, 2019), and correlations of ".33 to .51 with self-reported measures of the same or similar constructs" (Arctic Shores, 2018). And it states, as of 2022 and about that product: "To date, there has been no peer-reviewed research of this particular assessment tool."
How far any of this transfers to the current product we do not know, and we found no public data that would settle it. It is also worth noting that the same author writes about the whole class of instruments, not only this one, that "currently, there is not enough evidence to suggest that GBA meets these guidelines" in the context of ITC requirements — a statement that applies to our own category as much as to anyone's.
That last point is not a rhetorical flourish. NeuroFrame has no peer-reviewed publication of its own validation studies either. If the absence of peer review is your decisive criterion, it disqualifies both products on this page, and the honest thing is to say so here rather than to bury it further down.
Regulatory status: who owes what, and to whom
Start with the fact most often stated backwards. Under New York City's Local Law 144, the duty to commission and publish a bias audit sits with the employer or employment agency, not with the assessment vendor: the DCWP rules address "an employer or employment agency" throughout §§ 5-301 and 5-303. A vendor publishing an audit does so voluntarily, as a maturity signal. So the absence of an audit on a vendor's own site does not by itself say anything about that vendor's compliance, and nobody should present it as evidence of a breach. We are not lawyers and this is not legal advice: how the duty falls in your own deployment is worth checking with counsel.
What Arctic Shores states on its Compliance page, checked 4 August 2026: compliance with GDPR and the Data Protection Act 2018, and ISO 27001 certification. We did not find, on that page, information about the company's position on the EU AI Act or about an independent bias audit. We did not run an exhaustive search of the whole site, and we make no claim about what exists elsewhere.
Separately, and consistent with how the law allocates the duty: FDM Group, an employer client, publishes an NYC applicant bias audit page that lists Arctic Shores among its AEDT suppliers and links to a results PDF. We were not able to verify the contents of that PDF — auditor, dates, impact ratios — so we report only that the page exists and makes that listing, as of 4 August 2026.
On the EU AI Act, one date on the market is now wrong more often than it is right. Annex III, point 4(a) does classify systems used to evaluate candidates as high-risk. But the obligations under Chapter III, Sections 1–3 no longer apply from 2 August 2026: Regulation (EU) 2026/1744 of 8 July 2026, published in the Official Journal on 24 July 2026, moved that date to 2 December 2027 for Annex III systems. Two things were not moved: the transparency obligations of Article 50 apply from 2 August 2026, and the Article 5(1)(f) prohibition on inferring emotions in the workplace has applied since 2 February 2025. The prohibition is aimed at inferring emotions; how it applies to a given assessment depends on what that assessment actually processes — behavioural telemetry such as clicks, timings and choices, or face and voice. Which of these any particular product uses, and how its vendor's counsel reads Article 5(1)(f), is a question for that vendor and for your own lawyers, not for us.
On complaints and proceedings: we found none involving Arctic Shores in open sources as of 4 August 2026. That is a statement about our search, not a conclusion about the vendor, and we would report a filing as a filing rather than as a finding of fault in any case.
Finally, the context that matters for both of us. Research presented at FAccT 2025 collected 44 published bias audit reports covering 116 audits, which the authors found for approximately 2% of Fortune 500 companies. Of those audits, 53% include at least one impact ratio below 0.8, and 83% report missing race and/or sex information for some applicants. And the four-fifths rule itself is an evidentiary heuristic, not a safe harbour: 29 CFR § 1607.4(D) says a ratio above four-fifths "generally will not be regarded" as evidence of adverse impact, and immediately adds that smaller differences may still constitute it where they are significant in both statistical and practical terms. A published audit is a starting point for a conversation, not the end of one.
Where NeuroFrame's approach differs
Three differences are structural rather than cosmetic: one continuous session instead of a battery of discrete tasks; an externally built requirement reference instead of a model trained on the client's own staff; and a human-readable justification behind every benchmark cell. The first is described here, the second in the next section but one.
The session runs 30–60 minutes as a single dynamic simulation with delayed consequences, and produces eight parameters — three in the cognition domain, five in personality. The reason for a simulation with delayed consequences is specific: this class of task measures complex problem solving, a construct that overlaps general cognitive ability by only about 18% of variance (r ≈ .43; Stadler et al., 2015, meta-analysis of 47 studies). In other words it contains something classical ability tests do not capture — which is an argument for adding it to a process, not for removing anything from one.
Length has a second, narrower rationale. The vigilance decrement — the decline in sustained attention as time on task grows — is one of the most replicable effects in cognitive psychology, at group level. A long session therefore lets you observe behaviour both fresh and fatigued. We deliberately do not turn that into an individual metric of "accumulated fatigue": difference scores have poor test–retest reliability, and a number we cannot defend is worse than no number.
Now the honest constraints on all of this, because they cut against us. A systematic review of 34 studies (Ramos-Villagrasa et al., 2022) concludes that the game format does not offer advantages sufficient to recommend it over conventional methods, except in candidate reactions — so the defensible claim is comparable validity with substantially better candidate experience, not superiority. And on the Sackett, Zhang, Berry & Lievens (2022) re-estimates, the structured interview stands at .42, ahead of work samples at .33 and assessment centres at .29 — high-fidelity simulation does not automatically win. Length is not a validity argument, and we do not make it one.
Our own published psychometrics, for symmetry: Cronbach's α 0.69–0.77, test–retest above 0.83, CFI 0.96 for the two-domain model, R² 0.46 and AUC 0.77 against real KPIs and manager ratings on a validation sample above 3,000. These are our own figures, on our own sample and against our own criteria. They are not comparable with the Sackett et al. re-estimates quoted above — those are meta-analytic correlations across many studies, samples and outcome measures — and placing the two side by side would be an arithmetic that does not hold. The α range sits below the conventional 0.80 threshold, and we publish it rather than round it.
A battery of tasks and a single session: what actually differs
One property of format is measurable without an opinion: how findable and rehearsable a task is before the candidate sits it. A discrete, named, repeatable exercise can be identified, written up and practised. A single continuous simulation whose consequences unfold over half an hour is harder to reduce to a walkthrough. That is a claim about shapes of instruments, not about any vendor's quality.
The measurable trace is search demand, and here the measurement is ours rather than anyone's publication. Per Semrush, export of 4 August 2026, search strings pairing the vendor's name with an individual task name total roughly 550 a month ("balance" 170, "order" 110, "predict" 110, "lock" 90, "ticket" 70), and "arctic shores practice test free" runs at 390. Treat those as our own figures from a paid keyword tool, not as something you can open and check on a public page. What the pages behind those searches contain, and how well they correspond to the assessment, we did not examine and do not claim.
What it does not establish is anything about outcomes at Arctic Shores specifically. Whether preparation moves scores on their assessment, and by how much, we do not know: we found no public data either way. The vendor's own published position on adversarial attempts is the ChatGPT experiment noted earlier, from 3 October 2023, which reported that the models tested could not complete the assessment.
Other products are published as a single continuous session too. Sova Assessment describes Immerse as a single, continuous project in three stages and asks candidates to "complete a 15–30 min role simulation with virtual colleagues and customers"; Capsim writes that "typical inbox simulations are completed in just 15-60 minutes"; Owiwi's product page frames the candidate's experience as one journey — "Your journey through the Isles of the Shroud has begun…" — and states no completion time that we found there (vendor pages, checked 4 August 2026 — see Sources). In our reading these reproduce a working situation rather than turning on strategic decisions whose consequences arrive later; that is a description of format, not a ranking. We make no claim about the rest of the market.
An external reference versus a model trained on your own staff
Every assessment eventually has to answer the question "suitable for what?", and there are two ways to answer it. Either you learn the answer from the client's current high performers, or you define it outside the client entirely and then compare. The two produce different failure modes, and the choice is worth making consciously.
NeuroFrame's requirement library is built externally: 287 professions across 8 stages of the Adizes corporate lifecycle, 2,296 cells in total, with the profession list drawn from the Russian Ministry of Labour occupational register and professional standards. Two consequences follow. It works on day one, with no requirement that the client already employ dozens of people in the role. And it does not structurally encode the existing composition of the workforce, because the workforce was never an input.
The lifecycle dimension is not decorative. Computed across the 2,296 records: for 99% of the 287 professions, the requirement ranges at "Infancy" and at "Bureaucracy" overlap by less than half, with a median range overlap of 0.40. Of the 153 pairs where the ranges do not intersect at all between the extreme stages, 152 are risk appetite. Measured as movement of the range centre across the lifecycle: risk appetite shifts 18.8 percentile points, openness 15.4, conscientiousness 11.5, cognition 9.0; result focus barely moves at all.
Explainability was designed in rather than added. Each of the 2,296 cells carries a human-readable justification — 2,296 unique texts, 1.07 million characters, an average of 464 characters per cell, of which 32% name the rule that fired outright. And the benchmark can express tension rather than a single sum: in 303 of the 2,296 cells (13%) it permits high risk appetite together with a low floor on cognition or progress monitoring — which is how a method says "fits on every parameter and still warrants attention". A single composite number carries different information than a per-parameter benchmark with tolerances: the composite answers "how close overall", the benchmark answers "close on what, and where the tension sits". Which of the two a process needs is a question for the buyer, not a ranking.
One structural consequence worth stating plainly: the client receives anonymised codes and distributes them itself. NeuroFrame does not physically receive a name, a gender or an age, so the model cannot learn on protected characteristics. This is a design property, not a substitute for a bias audit — see the next section.
For contrast, what Arctic Shores publishes about norming: results are compared against norm groups, and the undated University Pack describing Skyrise City records the practice as "Employers tend to deselect below the 50th percentile on our fit scores", with decisions taken on the composite fit score rather than on individual traits. We have no current vendor document on the norming model beyond that, and we did not find one.
Questions worth putting to both vendors
The most useful outcome of a comparison page is not a choice but a better shortlist of questions. These are the ones we would ask, and they are deliberately symmetrical — every one of them is uncomfortable for us too.
Ask for the technical report, not the science page: reliability coefficients, validity coefficients, sample sizes, dates. Ask what those samples looked like and whether any of them resemble your candidate population. Both vendors on this page will have to answer that request rather than point at a website — and the answer you get is itself informative.
Ask for adverse impact figures computed on your own candidates, not on a pooled corpus. Under the New York rules, an audit built on pooled data across employers is permitted, which means a published summary may contain very little of any one employer's actual hiring. Ask, specifically, how many candidates fell into an unknown demographic category. This is not a hypothetical concern: the FAccT 2025 analysis of published Local Law 144 audits found that 83% of audits report missing race and/or sex information for some applicants, and that under the demonstrative assumption that all missing candidates belong to a single group, 70% of impact ratios would have a possible lower bound below 0.8. The authors state that assumption is unlikely to hold in practice, and report that under a proportional one the share falls to 28% — we quote both, because quoting only the first would be the kind of arithmetic we object to elsewhere on this page.
Ask how much weight the score carries in the decision. This is the operative test in the New York rules — a simplified output that outweighs any other single criterion, or overrides a human conclusion, brings the tool inside the law regardless of how the vendor labels it. The answer is a property of your process, not of the software, and a vendor cannot answer it for you.
Ask what a rejected candidate receives. GDPR Article 22(3) requires at minimum the right to human intervention, to express a point of view and to contest the decision; the Court of Justice of the EU, in Dun & Bradstreet Austria (C-203/22, judgment of 27 February 2025), addressed how far commercial confidentiality can limit that explanation; per the Court's own press release, a national rule may not as a general matter exclude the right of access where trade secrets are involved, and the balancing has to be done case by case. Read the judgment itself before relying on this in a procurement document. Then ask about the region: which norm groups apply to your market, in which languages the candidate experience and the report exist, and what the documented policy is for candidates who need adjustments.
When to choose Arctic Shores rather than us
Choose Arctic Shores if you need a vendor with a track record in UK and European volume hiring and a public client portfolio you can reference. Their site carries 16 case studies as of 4 August 2026, with named clients including Siemens, Capita, Adecco, Kantar, Thales and WSP. NeuroFrame has nothing comparable in public. If your board wants to see who else went first, that argument is not on our side.
Choose Arctic Shores if your process sits under GDPR, the EU AI Act or New York's LL 144 and you need a compliance package now. Our compliance perimeter is Russian: 152-FZ is closed architecturally, but we have no GDPR, EU AI Act or NYC LL 144 package yet. Arctic Shores states GDPR, the Data Protection Act 2018 and ISO 27001 certification on its Compliance page, and an employer client (FDM Group) publishes a bias audit page listing them as an AEDT supplier. On paper, that is further along than we are.
Choose Arctic Shores — or anyone else — if a formal bias audit and an adverse impact ratio report are decision criteria. NeuroFrame has neither. Our architecture means we never receive a name, gender or age, so the model cannot train on protected characteristics; that is a real property and it is still not an audit. If someone tells you it is, they are selling.
Choose a battery if you need coverage we do not offer. We run one task, which means we do not measure verbal or numerical ability at all. If your role screen needs those, we are not a substitute for a battery — at most an addition to one. Similarly on volume economics: client candidate material reports Arctic Shores windows starting at 15 minutes (Teach First: "anywhere between 15-50 minutes", checked 4 August 2026), against our fixed 30–60. On a wave of tens of thousands of applications, where every candidate minute costs completion, a configurable and potentially much shorter format wins on throughput, and ours is neither configurable nor short.
Choose someone else if your hiring is senior. 83% of our reference library is specialist-level roles; top management is covered by nine professions. And if you need ATS connectors, we have none — a direct consequence of the anonymised-code architecture, and a real operational cost you would carry.
Three more honest constraints. Our psychometrics are published but modest by the standards of long-established instruments: internal consistency α 0.69–0.77 sits below the conventional 0.80 threshold, and the comparison sample of 14,850 people is smaller than what forty-year-old instruments carry. And our client-facing profile handbook exists only in Russian; if your process requires Arabic materials, ask us about it directly before assuming. On adjustments for candidates with disabilities, our policy is not yet formalised; Arctic Shores' brochure describes its coverage as "Reasonably adjusted for most physical impairments and disabilities", and the details are worth confirming with that vendor directly.
Questions
- Is Arctic Shores a game or not?
- The vendor's materials use both words as of 4 August 2026. The site promotes the product as a task-based assessment — short interactive tasks with no direct questions — while the glossary section of the same domain describes the solution as "our game-based assessment". The historical flagship, Skyrise City, is described in the vendor's University Pack as running across nine levels. We record both formulations and do not interpret why they coexist.
- How many tasks are there, and how long does it take?
- There is no fixed number: client candidate material states that composition varies with the role. The brochure describes "up to eight engaging tasks"; the candidate guidance page says "up to nine tasks". On duration the vendor publishes "Know true potential, in 45 minutes", while Teach First asks candidates to "set aside anywhere between 15-50 minutes to complete the assessment" in its own candidate material (checked 4 August 2026). For comparison, a NeuroFrame session runs 30–60 minutes and is not split into separate tasks.
- Does Arctic Shores publish reliability and validity coefficients?
- On the four public pages we checked on 4 August 2026, no numerical coefficients and no sample sizes appear. That does not mean such data does not exist: the vendor's technical manuals (Skyrise City Technical Manual, 2019; Cosmic Cadet Technical Manual, 2018) are cited in the academic literature — we simply did not find them in open access. For current psychometric data, ask the vendor directly; in procurement that is a reasonable request.
- Does Arctic Shores have a bias audit under the New York law?
- Start with who owes what: under LL 144 the duty to commission and publish an audit sits with the employer, not the assessment supplier — that is how §§ 5-301 and 5-303 of the DCWP rules are written. Employer client FDM Group publishes a bias audit page listing Arctic Shores among its AEDT suppliers; we could not verify the contents of the linked report. The vendor's own Compliance page names GDPR, the Data Protection Act 2018 and ISO 27001 as of 4 August 2026; we found no audit information there. NeuroFrame has no formal bias audit at all.
- Does Arctic Shores work in the UAE, and in Arabic?
- We found no confirmation of an Arabic interface localisation in open sources as of 4 August 2026. In the current case-study list on the site (16 entries) we found no Middle Eastern clients — the geography shown there is the UK and Europe, including Greece — and the KPMG UAE case page referenced by earlier materials returns 404. The reason is not known to us and could be a site restructure as easily as a change of client roster. The vendor states that its assessment measures behaviour in tasks rather than answers to questions (How it works, checked 4 August 2026). How much text a candidate actually meets in the tasks we did not verify, so we draw no conclusion about the language barrier either way. Norm groups, in any case, are built on demographics and need to be local.
- What actually makes NeuroFrame different?
- Three things. One continuous simulation instead of a battery of discrete tasks: a dynamic problem with delayed consequences measures complex problem solving, a construct that overlaps general cognitive ability by only about 18% of variance. An externally built requirement benchmark instead of a model trained on your own staff: 287 professions × 8 lifecycle stages = 2,296 cells, drawn from an occupational standards register rather than from your high performers. And explainability by construction: every one of the 2,296 cells carries a human-readable justification, and 32% of them name the rule that fired outright.
- Can candidates prepare for either assessment in advance?
- Preparation is easier to organise around a discrete, named task than around a single long session — that is a property of format, not of quality. The measurable trace is search demand, and the measurement is our own: per Semrush, export of 4 August 2026, search strings pairing the vendor's name with an individual task name total roughly 550 a month, and "arctic shores practice test free" runs at 390. Those come from a paid keyword tool, not from a page a reader can open. Whether preparation moves final scores at Arctic Shores, and by how much, we do not know: we found no public data either way. The vendor itself published an experiment on 3 October 2023 reporting that the language models tested, and a home-built OCR bot, could not complete the assessment.
Sources
Every link was opened and checked on the date shown above.
- Arctic Shores — homepage — Arctic Shores
- sitemap.xml (lastmod dates for insights pages) — Arctic Shores
- ARCTIC SHORES LIMITED — company record — Companies House (UK)
- How it works — Arctic Shores
- Support For Candidates — Arctic Shores
- The Science Behind Our Assessments — Arctic Shores
- Compliance — Arctic Shores
- Glossary: What are Game Based Assessments? — Arctic Shores
- Case studies archive — Arctic Shores
- Arctic Shores closes £5.75m Series B — Arctic Shores
- Arctic Shores appoints Estelle McCartney as CEO — Arctic Shores
- Arctic Shores launches major update: Skill-enablers — Arctic Shores
- Introducing the world's first task-based assessment for Learning Agility — Arctic Shores
- ChatGPT vs Task-based Assessments: Can we break our own assessment? — Arctic Shores
- See more in people, with Arctic Shores (vendor brochure, undated, references place it circa 2021–2022) — Arctic Shores
- Arctic Shores University Pack (vendor PDF, hosted by the University of Southampton; describes Skyrise City) — Arctic Shores / University of Southampton
- Task-based assessment (client candidate material) — Teach First
- Sova Immerse — product page ("a 15–30 min role simulation", three stages of a single continuous project) — Sova Assessment
- CapsimInbox — Inbox Simulations ("typical inbox simulations are completed in just 15-60 minutes") — Capsim
- Owiwi — Product (single journey through "the Isles of the Shroud"; no completion time found on the page) — Owiwi
- Close, L. K. (2022). The psychometric structure of a game-based assessment. PhD thesis — Birkbeck, University of London
- Close, L. K. (2022) — full text (PDF) — Birkbeck, University of London
- NYC Applicant Bias Audit — FDM Group
- Notice of Adoption of Final Rule: Use of Automated Employment Decisionmaking Tools (6 RCNY §§ 5-300–5-304) — NYC Department of Consumer and Worker Protection
- Regulation (EU) 2026/1744 of 8 July 2026 (Digital Omnibus on AI), OJ 24 July 2026 — Official Journal of the European Union
- EU AI Act, Annex III (high-risk use cases, point 4(a)) — artificialintelligenceact.eu
- Red lines under the EU AI Act: the prohibition of emotion recognition in the workplace — Future of Privacy Forum
- Article 22 GDPR — Automated individual decision-making, including profiling — gdpr-info.eu
- Press release No 22/25, Case C-203/22 Dun & Bradstreet Austria (27 February 2025) — Court of Justice of the European Union
- 29 CFR § 1607.4 — Information on impact (Uniform Guidelines on Employee Selection Procedures, 1978) — Cornell Legal Information Institute
- Gerchick et al., Auditing the Audits: Lessons for Algorithmic Accountability from Local Law 144's Bias Audits, FAccT '25 — ACM FAccT 2025
- Sackett, Zhang, Berry & Lievens (2022). Revisiting meta-analytic estimates of validity in personnel selection, Journal of Applied Psychology 107(11), 2040–2068 — American Psychological Association
- Stadler, Becker, Gödker, Leutner & Greiff (2015). Complex problem solving and intelligence: A meta-analysis — Intelligence 53, 92–101
- Ramos-Villagrasa, Fernández-del-Río & Barrada (2022). Does evidence support gamification in personnel selection? A systematic review — Journal of Work and Organizational Psychology 38(1), 39–48