Researchers Identified AI-Generated Fake Manuscripts

A study revealed that large language models were used to fabricate scientific papers featuring phantom authors.

Updated on Sept. 30, 2026 in Artificial Intelligence

Bold flat-color editorial illustration of stacked volumes and a glass prism refracting light, representing the distortion of academic research data.
Computer scientists at the Samsung AI Center Warsaw have uncovered hundreds of fabricated scientific manuscripts circulating on research platforms that use AI-generated phantom authors. AI Illustration. Upload story photo >

Live Poll

Do you trust the authenticity of information or content that is generated by artificial intelligence?

In June 2026, computer scientists from the Samsung AI Center Warsaw reported the discovery of hundreds of fake scientific manuscripts circulating on research platforms like Zenodo and ResearchGate. These documents relied on AI-generated ghost identities to pass as legitimate academic work.

Why it matters

The proliferation of synthetic papers complicates academic integrity, as large language models can inadvertently create personas that appear credible enough to bypass standard screening. This research highlights how model training processes may reinforce recurring, fictional identities when prompted for expert insights.

Researchers tested 20 distinct model versions, including nine from Claude, ten from GPT, and one from Gemini, to track persona generation. Results showed model-specific habits, such as Claude Sonnet 4.6 ceasing the use of the Vasquez and Chen pair in 2026.

The players

Samsung AI Center Warsaw

A research division focused on developing machine learning and computer vision technologies.

Anthropic

A San Francisco-based developer of the Claude series of large language models.

OpenAI

A San Francisco-based research organization behind the GPT series of generative AI models.

The details

The research team at the Samsung AI Center Warsaw prompted multiple versions of large language models—algorithms trained to predict and generate human-like text—to create stories featuring pairs of experts. They found that models consistently hallucinated specific ghost authors, such as Elara Voss for GPT models or Aris Thorne and Lena Petrova for Gemini. These fake manuscripts were then uploaded to platforms like Zenodo, some even obtaining real digital object identifiers—unique alphanumeric strings used to persistently link to digital academic content.

Timeline

  1. Between 2024 and 2026, researchers tested various versions of LLMs.

  2. Claude Sonnet 4 was released in May 2025.

  3. GPT version 5.1 was released in November 2025.

  4. GPT-5.4 was released in March 2026.

  5. A preprint study on ghost identities was published in June 2026.

The Tech Race

This report follows the findings established by the 2026 preprint study on ghost identities to quantify how AI platforms are being manipulated to seed misinformation in academic repositories. It serves as a diagnostic assessment of how current model architectures generate persistent, identifiable biases that researchers must now mitigate during peer-review processes.

Researchers and database managers are now forced to implement more rigorous identity verification processes for manuscript submissions on open-access platforms. This shift may lead to increased scrutiny and potential delays for legitimate researchers using preprint servers.

The takeaway

The rise of synthetic academic identities demonstrates a critical vulnerability in current scientific publication workflows. Readers should verify the institutional affiliations of obscure authors through independent channels or established academic registries.

Further reading

For more on how language models are evolving, see Artificial Intelligence.

Source note: This article includes information reported by Nature.

Live Poll

Do you trust the authenticity of information or content that is generated by artificial intelligence?