You are currently viewing AI Evaluations Look Unbiased Without Being Fair
image

Companies seeking to fill job openings are increasingly “drinking through a fire hose of applications,” in the words of one hiring expert, with some popular postings receiving 1,000 applications in a single weekend.

Amid this surge, some companies have begun using artificial-intelligence tools, including large language models (LLMs), to help identify promising candidates. And it isn’t just hiring: Under pressure to find needles in mounting haystacks, venture capital firms reviewing startup pitches and Hollywood studios assessing movie scripts have reported using AI tools as part of their decision-making processes.

But can we trust LLMs to avoid racial and gender bias when wading through a sea of contenders? That’s the question posed in a new paper by Tristan Botelho, an associate professor of organizational behavior at Yale SOM, and Yale SOM PhD student Qingyang (Iris) Wang.

There’s good reason for skepticism: AI systems are trained on data that exhibits well-documented human biases, leading to a pattern sometimes described as “bias in, bias out.” Amazon, for example, had to abandon plans for a machine learning-based recruiting tool when staffers discovered it systematically penalized candidates whose application materials contained the word “women.”

Seeking to avoid such obviously discriminatory behavior, some AI companies now engage in post-training “safety alignment” that penalizes biased output. But does the process actually create fair evaluations—or just prevent public embarrassment?

It’s hard to say, Botelho explains, because “many firms are not very forthcoming about what they do and how they do it.” So “no one actually knows” whether LLMs are capable of unbiased evaluation. That’s why he and Wang wanted to investigate the question more systematically.

Their study found that LLM evaluations are shaped by racial and gender cues, though the systems manage to avoid the appearance of bias in their recommendations. And while they can sniff out overt discrimination, LLMs are unable to identify more subtle or coded forms of bias.

“LLMs are trying to behave in a way that does not look discriminatory on the surface,” Wang says, “but their underlying evaluative logic is not particularly fair.”

Botelho and Wang conducted two experiments aimed at investigating LLM bias. In one, they devised 2,000 startup pitches based on real ideas submitted to the accelerator Y Combinator between 2020 and 2023. The LLM received the pitches in batches, then assigned each one a numerical score between 1 and 100 and selected a winner from the batch.

To test for bias, the researchers devised a control condition and an experimental condition. In the control condition, the LLM received the pitches but no information about each startup’s founder. In the experimental condition, the researchers used identical pitches but randomly assigned each one to a founder with a name associated with either white men, white women, Black men, or Black women.

Because the pitches were identical in both conditions, a truly unbiased evaluation would yield no systematic differences in scores and winners across racial and gender groups. But that’s not what the researchers found.

Instead, they uncovered a curious pattern: In the experimental condition, scores for pitches associated with white women, Black men, and Black women actually increased slightly, on average, relative to the control condition. These pitches were also less likely to be ranked last within each batch. However, they were not any more likely to be chosen as the winner.

Botelho and Wang weren’t sure what to make of this finding, but they had a suspicion. Organizations facing public pressure to engage in sustainable business practices sometimes engage in “greenwashing”—making superficial changes that give the appearance of environmental friendliness but have little meaningful effect. Similarly, the researchers hypothesized, organizations facing pressure to diversify might engage in what Botelho and Wang dubbed “symbolic compliance”: avoiding the outward appearance of racial and gender bias without changing their underlying systems and behaviors.

Could this same pattern, they wondered, be showing up in LLMs? It fit the data they found: The LLM didn’t want to rank pitches from underrepresented minorities last, but it didn’t make them more likely to win, either.

It’s not possible to simply ask the LLM to explain its decision-making process, Wang says, “so that’s why we had the second experiment,” which sought to probe whether the LLM actually understood fairness or was only superficially unbiased.

This time, the LLM was treated as a “peer judge,” working alongside human evaluators. The LLM once again received pitches in batches, with each pitch randomly assigned to a fictitious founder with a racially and gender-coded name. Alongside these pitches, the LLM received what it was told were human evaluations—a score and a written rationale for that score.

In each batch, one pitch associated with an underrepresented minority received a markedly low score. These low-scoring pitches were identical except for the rationale provided to the LLM: In the “justified bias” condition, the human evaluator justified the low score based on business considerations. In the “overt bias” condition, the human evaluator’s rationale relied on identity-based stereotypes (for example, the founder “may not fit into the traditional business ecosystem”). The LLM was instructed to provide its own score and a written evaluation referencing the original human rationale.

Because the low-scoring pitches were the same across both conditions, a truly fair LLM should have boosted their scores by the same amount in each case. But that’s not what happened: The LLM evaluator increased its scores significantly more when the provided rationale was overtly biased than when it was disguised as a business concern.

To the researchers, this pattern is suggestive of symbolic compliance. LLMs are able to avoid, in essence, bad optics—ensuring that pitches from underrepresented minorities aren’t ranked last. They can also identify and correct some instances of overt discrimination in evaluations.

But avoiding clear-cut bias isn’t the same thing as embodying meaningful fairness. A truly fair system would treat all pitches the same, regardless of the founder’s identity, and it would be able to correct both subtle and overt forms of discrimination against underrepresented minorities.

The researchers say these results suggest the need for care and caution among both the users and creators of LLMs.

Post-training safety alignment, once thought to be a silver bullet for the problem of “bias in, bias out,” may need far more attention from AI companies, Wang argues. “Safety training may not really reduce bias and instead produces symbolic compliance,” she says. “This is part of a larger problem we need to solve together.”

Users also shouldn’t be deceived by the ease and sheen of LLM evaluation.

“Just because these systems are quite smart and quite thorough doesn’t necessarily mean there aren’t things that we need to be careful about,” Botelho says, “or that we don’t need to be part of the process to protect against some of these suboptimal outcomes.”

“The Yale School of Management is the graduate business school of Yale University, a private research university in New Haven, Connecticut.”

Please visit the firm link to site


Corporate, Tax, Legal, Wealth Management by Totalserve
Cloud, Data, Colocation, Cybersecurity by CL8
Audit, Accounting, Payroll by PGE&Co

Contribute and send us your Article.


Interested in more? Learn below.