George Washington University

07/21/2026 | News release | Distributed by Public on 07/21/2026 09:01

Three GW Graduates Publish Paper on Detecting Bias in GPT Models

Three GW Graduates Publish Paper on Detecting Bias in GPT Models

The team of GW alumni developed a framework that can be used by financial institutions to help safeguard against biased and unethical decisions.
July 21, 2026

Authored by:

Brook Endale

From left: Hanna Courtot, Patrick Hall, Shegufta Tasneem and Katherine Fullowan.

As institutions increasingly rely on artificial intelligence to help inform decisions that can impact people's financial futures, employment opportunities and access to services, experts have been raising concerns on whether these systems treat people fairly and produce unbiased outputs.

In response, researchers and industry practitioners have been working to develop methods to evaluate AI systems and identify whether they reinforce systemic biases that could lead to unethical or unlawful decisions.

Three recent graduates of the George Washington University School of Business' Master of Science in Business Analytics program contributed to that effort by developing a framework to evaluate LLMs as part of their practicum capstone project.

"Companies are rushing to incorporate AI, generative AI and large language models into their current system. And because of the rush, there is a very timely need to test LLMs before they are implemented, before any repercussions happen," said Shegufta Tasneem, M.S. '25, M.B.A. '25, who, along with fellow alumni Hanna Courtot, M.S. '25, M.B.A. '25, and Katherine Fullowan, M.S. '25, developed the framework while working with consulting firm EY.

Their research, which was later published in the peer-reviewed journal "Mathematics," explores how organizations can test whether AI-generated responses differ based on demographic characteristics.

As organizations rapidly adopt AI tools, the graduates said it's important to not assume they are always objective.

"Companies need to be aware of the output and the potential bias that could show up," Courtot said. "Blindly trusting LLMs and just relying on the first output that you see in the spirit of moving fast and doing the 1% better is nice conceptually, but you have to think of the impact that it's having on all the people who will interact with its output."

She said the stakes are especially high when it comes to financial decisions like lending practices.

The team was able to design a multi-step framework to help financial institutions evaluate LLMs. The framework tests whether a model produces different responses when details such as a person's race or gender are changed while the rest of the prompt remains the same.

Rather than relying on a single measurement, the framework combines several evaluation methods, alongside human review to identify subtle shifts and differences in the outputs.

The students tested the framework using news stories about financial fraud, asking multiple GPT models to summarize the same stories while changing only the demographic identity of the individual in the story.

Then, they analyzed whether the summaries differed in wording, tone, sentiment and overall similarity. While the differences were often small, the framework was able to detect statistically significant variations across some demographic groups.

"There are so many forms of bias and it's not just one simple definition," Courtot said. "So that's why our approach is so layered, and has so many pieces. So that it is really robust and it's trying to capture bias in any way that it can show up."

One of the challenges, they added, was even defining what bias looks like in a way that can be measured consistently across different scenarios.

"Identifying bias in and of itself is something that doesn't really have one specific definition," Fullowan said. "Based on different use cases, you could define it differently, especially in our case, we were looking at bodies of text. So to quantify whether or not something is biased in a reproducible manner was something we spent a lot of time kind of trying to workshop."

The team also noted that part of what makes their framework useful is that it can be applied across different models and updated as systems evolve. While testing multiple GPT models, the team found that newer versions generally showed fewer statistically significant differences across demographic groups, indicating that efforts to refine safeguards on AI models may be working.

"We tested different GPT models and saw how results progressed across versions," Tasneem said. "It wasn't drastically different, but there were clear changes. That validated our framework because that's what we would have expected to see, right? Because these companies are putting guardrails on the models. We would expect the results to get better. And that's exactly what we saw."

While newer models generally performed better, Courtot said that doesn't necessarily mean institutions aren't continuing to rely on older versions. "Some older models are still very common because they're more affordable, and we still saw bias in those outputs," she said.

Looking ahead, the students hope their work contributes to encouraging organizations to take a more thoughtful approach to adopting AI by evaluating these systems before putting them into use.

"Companies are moving so fast that sometimes they're not seeing the impacts of rolling these models out," Courtot said. "We're trying to give them a way to pause and actually measure what's happening before it reaches users."

The team said while their work is a starting point, they see it as something that could be expanded and refined further.

"I think what we built is very baseline," Tasneem said. "It would be great to see it implemented on a larger dataset or used internally by a company to see how it performs in practice."

While working on the project, the students were also taking a responsible machine learning course taught by Patrick Hall, a teaching assistant professor of decision sciences and the chief AI officer at the School of Business. They were able to apply lessons from the course to their practicum research.

Hall also served as a faculty mentor throughout the project, offering feedback on the students' methodology and helping guide them through the publication process.

"They did really good work on measuring potential algorithmic bias in language model outputs, which is a difficult problem," Hall said. "It's a less established question, and I thought they came up with a really solid way to measure that."

Hall said the students' work contributes to the rapidly evolving field of evaluating AI, where researchers are still developing reliable ways to assess whether LLMs produce fair and consistent results.

"In general, we rely more and more on automation, and we want these systems that are making decisions or helping us make decisions to be right and to be fair," Hall said. "There's really no guarantee of either of those things."

George Washington University published this content on July 21, 2026, and is solely responsible for the information contained herein. Distributed via Public Technologies (PUBT), unedited and unaltered, on July 21, 2026 at 15:01 UTC. If you believe the information included in the content is inaccurate or outdated and requires editing or removal, please contact us at [email protected]