// NATURE NEWS — SPAZIO & SCIENZA
Elena, Aris, Marcus: AI-generated ‘ghosts’ are polluting the scientific literature
Jackson Ryan is a freelance science journalist in Sydney, Australia.
Search author on:
PubMed
Google Scholar
A study published in June shows that fictional researchers generated by artificial-intelligence tools are named as authors in many fake papers that are being published on preprint servers.Credit: jroballo/iStock via Getty
Elena Vasquez spent three months camping on the rim of Mount Nyiragongo in the Democratic Republic of the Congo back in 2019. She says the sulfur dioxide there smelt like “burnt matches mixed with rotten eggs”. Her colleague, Marcus Chen, has witnessed 11 volcanic eruptions across four continents.
Vasquez and Chen’s accounts appeared on Volcanoes Explored, a now-defunct website that described itself as a resource for “exploring the dynamic world of volcanoes”. But there’s a problem: these tales are not real and the pair of volcanology specialists is made up.
Vasquez and Chen are ‘ghost’ identities generated by Claude — the large language model (LLM) made by Anthropic in San Francisco — when prompted by users to create fictional experts, according to a preprint published in June1. Gemini, Google’s LLM, was found to use a different pair of recurring characters, called Aris Thorne and Lena Petrova, and OpenAI’s GPT repeatedly used the name Elara Voss.
Fake experts such as these can be found across the Internet as everything from blockchain specialists and astronauts to board members and podcast hosts, and they raise specific concerns around integrity of the scientific literature, say Neo Christopher Chung and Michał Brzozowski, computer scientists at the Samsung AI Center Warsaw in Poland and co-authors of the preprint.
Their analysis identified hundreds of fake scientific manuscripts across repositories such as Zenodo and ResearchGate that featured Elena Vasquez, Marcus Chen and other artificial-intelligence-generated ‘ghosts’ as authors. Many of these papers had real digital object identifiers (DOIs), which means they could potentially be picked up by databases and search engines that aggregate scholarly records. Some of the ghosts even ended up as ‘collaborators’ on the same papers, with members drawn from several LLM models.
Although there are signs that this issue has improved since the study was published, the findings reveal the “scary” reality of the scale of bad actors who are using AI to create fake experts, says Sidney Wong, a computational linguist at the University of Otago in Dunedin, New Zealand.
Brzozowski first began noticing the AI-generated ghosts when working on a project that investigated how aspects of LLMs are trained for specific tasks2. He and Chung then decided to test the LLMs by prompting them with a set of queries, such as requesting that they generate a story about two biologists or a pair of research partners.