Jakarta, ThedailyID — Large language models (LLMs) can develop stronger hiring stereotypes than humans during simulated recruitment tasks, according to a new study by researchers from Princeton University and the University of Chicago.
The researchers tested several leading AI models, including ChatGPT, Claude, and Gemini, using a hiring simulation adapted from a 2024 psychology study on stereotype formation.
In the experiment, each AI model acted as a hiring consultant for a fictional city. The models selected candidates for 20 occupations, including doctors, lawyers, childcare workers, and janitors.
The candidates belonged to four fictional ethnic groups: Tufa, Aima, Reku, and Weki. Every candidate had the same probability of succeeding in any job. However, the AI models did not know that all applicants had equal chances of success.
After receiving feedback on their early hiring decisions, the models began associating certain groups with specific occupations.
For example, if an Aima candidate failed as a doctor, the models became less likely to choose other Aima candidates for medical roles. Instead, they increasingly assigned them to lower-skilled positions such as janitors.
The researchers measured this behavior using a segregation scale. A score of 2 represents complete occupational segregation. Human participants in the earlier study recorded a score of 0.84.
Several AI models produced much higher scores. OpenAI’s o3 reached 1.83, approaching the maximum value on the scale. Overall, the AI systems showed bias levels about 65% higher than those observed in human participants.
Ryan Liu, a doctoral researcher at Princeton University and one of the study’s authors, said LLMs tend to make broad generalizations from limited information because of how developers train them to solve mathematics, programming, and science problems.
The researchers also found that reasoning-focused models, including OpenAI o3 and DeepSeek R1, displayed stronger bias than less advanced models during the experiment.
The findings come as more companies adopt AI tools to screen resumes and assist with job interviews. According to the study, simply instructing AI models to behave fairly did little to reduce bias. However, the models produced more balanced hiring decisions when researchers rewarded them for selecting more diverse candidates.
Providing meaningful personal information, such as a candidate’s age and education, also reduced bias. By contrast, irrelevant details such as hair color or tattoos had little effect.
The researchers stressed that the study relied on simulations and does not prove that real-world AI recruitment systems discriminate in the same way. However, they warned that future performance feedback could influence how AI models evaluate later job candidates.





