RSF - Reporters sans frontières

10/08/2026 | Press release | Distributed by Public on 10/08/2026 10:36

How Chinese propaganda invades Western chatbots

Xi Jinping's narratives on China extend beyond national borders and have found their way into the latest versions of certain Western chatbots, such as ChatGPT or Claude. Reporters Without Borders (RSF) spoke to American researcher Hannah Waight, from the University of Oregon, who led a study on the subject.

Hannah Waight is an Assistant Professor of Sociology at the University of Oregon who researches media, information and authoritarian politics. In her interview with RSF, she shared insights from a recent study co-published in the science journal Nature on how state-controlled media can shape large language models.

Your study showed that Chinese propaganda has infiltrated most recent Western large language models. How is that possible?

We think this is a really interesting case of how state-controlled media can end up beyond the bounds of where it started. There has been a lot of focus on two mechanisms through which powerful institutions can shape large language models. One is regulation. When DeepSeek [a Chinese chatbot developed by the eponymous company] came out, for example, there was a lot of focus on how it might be parroting the state's voice, because China has some degree of regulatory control over the model. Another related concern is how institutions can shape these models through ownership - the example here being Grok [the chatbot embedded into the social media platform X]. People are very concerned that Elon Musk, who owns the company that created this model, might be exerting an ideological influence on it. These processes are important and deserve academic and media attention. But what we're really studying is a third mechanism by which powerful institutions shape the output of these models, and that is through the information environment itself.

China has put a lot of effort into shaping what's on the Chinese internet, through processes of propaganda and censorship that promote certain content. Since the mid-to-late 2000s, there have been efforts to archive what's on the web - through initiatives like Common Crawl, a non-profit organisation that crawls the web to provide open-source databases of web content - that have essentially hoovered up all this data, including data from the Chinese internet. And since around the late 2010s, that data has been repurposed for large language model training. In our paper, we argue that the inclusion of state-coordinated media from China in the training data for large language models is having an effect on these models, leading to more pro-PRC [People's Republic of China] model outputs [the content produced by an AI], especially when you query [give instructions to] these models in Chinese rather than English or other linguistically distinct comparison languages.

What happened is largely a story of unintended consequences: the decision of China's state-coordinated media to intervene in the information environment, then the decisions of the organisations that created these archives, and then - most crucially - the decision of companies like OpenAI and Anthropic to take all this data from the internet, treat it as a neutral repository for language and knowledge, and train their models on it.

So, this is more a side effect than a deliberate strategy from Chinese authorities?

We have no evidence that this is intentional on the part of the Chinese state. It wouldn't even make logical sense, because much of the training data we looked at was collected before these models existed. That said, our story does highlight a worrying possibility: now that this pathway exists, it changes the incentives of authoritarian regimes. It gives them another reason to coordinate their information environment because, by doing so, they can shape not only what their own citizens perceive, but the models that are now playing an increasingly important role in our broader information environment.

That said, this story is not just about China. In the final study of the paper, we provide evidence that this is very likely occurring in any authoritarian regime that puts significant effort into media control and also has a high degree of language exclusivity [e.g the country's main language is spoken almost exclusively within its territory]. One of the reasons China's system can have this effect is that there are very few other powerful organisations exercising this kind of control over information in the Chinese language. And that's true for a number of authoritarian regimes - like Vietnam and Turkmenistan. These are places where we believe the same process is similarly occurring.

What can be done to counter the spread of propaganda within large language models?

This is a really important question, and it doesn't have direct, easy answers. I want to come back to what I said at the start: the information environment mechanism is one of several ways that powerful institutions can shape these models. Another is regulation. One concern I, personally, have is whether regulation might end up being used for political ends, particularly in the United States. So we have to weigh these different risks carefully.

That said, as an author team, we believe a really important first step is transparency. We need more information about what these models are trained on, what's in the training data, and how it changes over time as companies make different decisions. That matters for consumers, so they can make informed choices about which models they use and what they use them for. It matters for policymakers, so they can understand what's actually inside these very complex systems. And it matters for the organisations that are increasingly deploying these models with sensitive data - they need to know the source of the information their systems rely on.

One key takeaway from our paper is that what these models effectively do is launder information, cutting it off from its source. We need to bring the source back in. Anyone using these models should be able to know where the information they're getting ultimately comes from. This ideal is much easier to state than implement in these complex systems, however.

Full study available here.

Image
178/ 180
Score : 13.85
Published on 08.10.2026
RSF - Reporters sans frontières published this content on October 08, 2026, and is solely responsible for the information contained herein. Distributed via Public Technologies (PUBT), unedited and unaltered, on October 08, 2026 at 16:36 UTC. If you believe the information included in the content is inaccurate or outdated and requires editing or removal, please contact us at [email protected]