AI hiring tools are repeatedly binning the same job seekers, with Black and Asian candidates taking a disproportionate kicking.
A Stanford-led study of 4mn job applications across 156 employers using the Pymetrics hiring platform found evidence of “systemic rejection” linked to its algorithms.
According to the Financial Times Pymetrics assesses applicants through a set of online games, which is the sort of thing HR departments like because it looks cheaper than reading CVs.
The research found job seekers would need to apply for at least 25 different roles to be almost certain of getting one recommendation to move to the next stage.
The study is billed as the largest examination of AI hiring algorithms so far. It adds to growing concern that automated recruitment tools are baking bias into hiring across large employers.
Game-based assessments have become popular with employers dealing with piles of applications. Firms such as Pymetrics and HireVue are increasingly used to sift candidates before a human being bothers looking.
Jobseekers have complained they spend hours completing tests with little chance that anyone with a pulse will read their application.
Northeastern University assistant professor of philosophy and computer science Kathleen Creel said:“As a single vendor comes to dominate decision-making in a space, their quirks or shortfalls can be present across that entire sector in a way that wasn’t possible before.”
The Stanford Institute for Human-Centred AI-led study examined 4 million job applications submitted through Pymetrics between December 2018 and December 2022.
The dataset covered 156 employers, most with annual revenues of $5bn or more.
Researchers found “clear racial disparities” in outcomes. Looking at individual roles, one in 10 positions in the dataset showed “adverse impact” against Black applicants. One in 20 roles showed an adverse impact on Asian applicants.
Adverse impact is used by US federal agencies to describe a selection rate for any race, sex or ethnic group that is less than four-fifths of the most selected group.
A previous study by University of California, Berkeley and University of Chicago researchers found that “distinctively Black names” cut the probability of employer contact by 2.1 percentage points compared with “distinctively white names”.
The Stanford-led research found several employers used identical algorithmic models to screen candidates for some roles. Researchers identified 42 models “shared across” different employers. That means candidates rejected by one company were likely to fail at others using the same model.
The data showed few candidates had been caught by that particular issue, which is not quite the ringing endorsement the HR software crowd might want.
Four per cent of applicants who applied for 10 roles were recommended for rejection across all of them by the platform’s algorithm. That was higher than chance would suggest.
Researchers wrote: “When applying to two positions at two different employers, applicants might reasonably expect that they are receiving two separate evaluations and therefore two chances. But if both positions share the same model, their numerical score will be identical.”
Pymetrics algorithms assess traits such as risk appetite and response speed, as well as characteristics such as trust and care for others. Applicants whose performance most closely matches that of top employees are recommended to advance. Everyone else gets the algorithmic thumbs down, usually without much ceremony.







