0% found this document useful (0 votes)
20 views5 pages

Nonsense Papers Persist in Science

Hundreds of computer-generated gibberish papers from SCIgen software are still appearing in scientific literature years after first being discovered. Researchers recently identified 243 nonsense papers from 2008-2020 across various publications. Many publishers are now investigating these papers and retracting them, with some journals having published over 50 gibberish papers each. While rare overall, these papers highlight issues with "publish or perish" cultures and lack of screening that allow nonsense to pass peer review.

Uploaded by

MarkAllenPascual
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
20 views5 pages

Nonsense Papers Persist in Science

Hundreds of computer-generated gibberish papers from SCIgen software are still appearing in scientific literature years after first being discovered. Researchers recently identified 243 nonsense papers from 2008-2020 across various publications. Many publishers are now investigating these papers and retracting them, with some journals having published over 50 gibberish papers each. While rare overall, these papers highlight issues with "publish or perish" cultures and lack of screening that allow nonsense to pass peer review.

Uploaded by

MarkAllenPascual
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd

Hundreds of gibberish papers still lurk in the

scientific literature
The nonsensical computer-generated articles, spotted years after the problem was first
seen, could lead to a wave of retractions.

 Richard Van Noorden

  
  

Nonsensical research papers generated by a computer program are still popping up in


the scientific literature many years after the problem was first seen, a study has
revealed1. Some publishers have told Nature they will take down the papers, which
could result in more than 200 retractions.

The issue began in 2005, when three PhD students created paper-generating software
called SCIgen for “maximum amusement”, and to show that some conferences would
accept meaningless papers. The program cobbles together words to generate research
articles with random titles, text and charts, easily spotted as gibberish by a human
reader. It is free to download, and anyone can use it.

By 2012, computer scientist Cyril Labbé had found 85 fake SCIgen papers in
conferences published by the Institute of Electrical and Electronic Engineers (IEEE);
he went on to find more than 120 fake SCIgen papers published by the IEEE and by
Springer2. It was unclear who had generated the papers or why. The articles were
subsequently retracted — or sometimes deleted — and Labbé released a
website allowing anyone to upload a manuscript and check whether it seems to be a
SCIgen invention. Springer also sponsored a PhD project to help spot SCIgen papers,
which resulted in free software called SciDetect. (Springer is now part of Springer
Nature; Nature’s news team is editorially independent of its publisher.)
The fight against fake-paper factories that churn out sham science

Labbé, who works at the University of Grenoble Alpes in France, originally searched
manuscripts for words typical of SCIgen’s vocabulary. But he and another computer
scientist, Guillaume Cabanac at the University of Toulouse, France, came up with a
new idea: searching for key grammatical phrases characteristic of SCIgen’s output.
Last May, he and Cabanac searched for such phrases in millions of papers indexed in
the Dimensions database.

After manually inspecting every hit, the researchers identified 243 nonsense articles
created entirely or partly by SCIgen, they report in a study published on 26 May 1.
These articles, published between 2008 and 2020, appeared in various journals,
conference proceedings and preprint sites, and were mostly in the computer-science
field. Some appeared in open-access journals; others were paywalled. Forty-six of
them had already been retracted or deleted from the websites where they were first
published.

Since last year, the researchers have added another 20 papers to their list, including
gibberish articles created by MATHgen (software that generates mathematics papers)
and the SBIR proposal generator (which creates nonsense grant proposals). Cabanac
and Labbé have posted some of their findings on Twitter and the post-publication peer
review website PubPeer, and they are releasing their full results online.

CV padding
Most of the latest batch of SCIgen papers were authored by researchers from China
(64%) or India (22%), although Labbé notes that the manuscripts could have been
submitted in anyone’s name without their knowledge. One author of several of the
papers told Labbé and Cabanac that he’d submitted them as hoaxes. But other
manuscripts appear to have been edited with genuine reference lists, suggesting that
they might have been generated to inflate scientists’ citation counts. “I think the vast
majority are created to pad CVs in order to fulfil a need to publish papers,” says
Labbé.

The researchers found only two SCIgen papers that hadn’t been retracted at IEEE —
which is evaluating both of them — and one Springer paper that included a fragment
of MATHgen text. But other publishers were caught out more badly. IOP Publishing,
a subsidiary of the London-based Institute of Physics, says it retracted ten papers “as
there was clear evidence they had been computer-generated” and is investigating why
they weren’t identified during peer review at the conference where they were
accepted. “We have reasonable evidence to suggest that the peer review process for
some of these papers was compromised,” says Kim Eggleton, the publisher’s integrity
and inclusion manager.

The publishers who posted the most SCIgen content were Trans Tech Publications, a
Swiss publisher, which published 57 SCIgen papers, Blue Eyes Intelligence
Engineering and Sciences Publication (BEIESP), based in India, which had 54; and
Atlantis Press, a French publisher that was acquired by Springer Nature this March,
with 39. Both Trans Tech Publications and Atlantis told Nature that they were
investigating and were in the process of retracting the articles, but a spokesperson for
BEIESP said that it published only articles with original content that passed double-
blind peer review and plagiarism checks.
Hundreds of extreme self-citing scientists revealed in new database

The popular SSRN preprint server, where papers are shared before peer review, had
published 16 SCIgen articles, the study found. A spokesperson for SSRN said it was
investigating the issue, and noted that it provided “limited screening” for its preprints
(with “advanced screening” for health-care manuscripts).

Cabanac is concerned by the non-transparent way in which some publishers deal with
such papers. The IEEE, for instance, has wiped some SCIgen papers off its website,
but left formal retraction notices for others. Cabanac also notes that research papers —
or earlier versions of them — sometimes disappear from the SSRN preprint server,
without such changes being recorded.

An IEEE spokesperson said that its policy on removing a paper or leaving a retraction
label was “contingent on the outcome of our evaluation”; SSRN did not respond to a
question about its policies on retraction or deletion.

SCIgen papers are extremely rare: Labbé and Cabanac estimate from their screen that
they make up a mere 75 papers per million in the computer-science literature. They
are a far smaller problem than are, for instance, suspected paper mills — which create
seemingly real research papers to order for academics — which Labbé and Cabanac
have also helped to uncover.

But, says Labbé, the existence of these papers is an indication of the harmful effects
of a ‘publish or perish’ culture, and an example of how nonsensical work can still
make it into conference proceedings or journals. “You shouldn’t find these things in
the literature,” he says.

Nature 594, 160-161 (2021)

doi: [Link]
UPDATES & CORRECTIONS
 Clarification 28 May 2021: This article has been updated to clarify a statement
from IOP Publishing.

Common questions

Powered by AI

The presence of SCIgen papers undermines the credibility of scientific publications by introducing nonsensical articles into the research community, which can lead to misinformation and erode trust in academic work. Publishers like IEEE and Springer have responded by retracting identified SCIgen articles and enhancing their screening processes, including funding projects like SciDetect to identify fake papers. However, responses have varied, with some publishers, like BEIESP, maintaining that their processes are robust despite evidence to the contrary, illustrating an uneven diligence in addressing the problem .

The effectiveness of transparency and retraction policies varies significantly across scientific journals. While some publishers like IEEE have taken steps to retract SCIgen papers, their inconsistent approach—such as removing papers without leaving formal retraction notifications—hampers transparency. Platforms like SSRN have faced criticism over a lack of clear retraction policies, affecting credibility and accountability. Effective retraction policies should include clear notices and maintain a transparent record for educational purposes and to prevent recurrence, ensuring systemic resilience against fraudulent content .

Labbé and Cabanac devised methods to detect SCIgen papers by initially searching for typical vocabulary and phrases characteristic of SCIgen-generated texts. They later refined their techniques to include searching for key grammatical phrases and using these findings in conjunction with databases to identify nonsensical articles. Their contributions are crucial as they provide tools and methodologies for publishers and researchers to combat the infiltration of fake science, thus helping maintain the integrity and reliability of scientific literature .

The SCIgen software, created by three PhD students in 2005, generated nonsensical but structurally plausible research papers to expose vulnerabilities in the academic publishing process. The ease with which these papers were accepted at conferences and journals without thorough peer review highlighted significant deficiencies in the vetting procedures of academic publishers. This incident revealed how some publishers lacked stringent checks to assess the content validity or relevance, relying instead on superficial or absent reviews that allowed gibberish papers to slip through .

Technological innovation facilitates both the creation and detection of fake research papers. SCIgen exemplifies how automated software can generate papers that appear structurally valid, exploiting weak spots in academically rigorous processes. Conversely, innovations like SciDetect and enhanced screening algorithms are pivotal in identifying these fake articles. These tools leverage technology to scrutinize submissions for typical fraud markers, thereby protecting the integrity of scientific literature from malicious or negligent activities .

SCIgen papers differ from suspected paper mills primarily in intent and production. SCIgen papers are generated using software to produce nonsensical academic articles to test the limits of peer review systems. In contrast, paper mills create seemingly legitimate, often plagiarized or minimally altered research articles to fulfill academic demands or profit motives. Both present challenges, but paper mills pose a greater threat because their outputs can appear credible and mislead academia, necessitating complex detection methods and widespread vigilance to safeguard academic integrity .

The IEEE has retracted SCIgen papers, sometimes without leaving formal notices, based on internal evaluations. SSRN has investigated the issue but has been criticized for its lack of transparency in policies regarding retraction and deletion of such papers. IOP Publishing discovered and retracted ten SCIgen papers, admitting a failure in their peer-review processes during conferences. These varying responses highlight different levels of effectiveness and transparency in how publishers and platforms address the problem, pointing to the need for standard protocols .

The existence of SCIgen papers underscores the pressures of the 'publish or perish' culture, where academics feel compelled to publish often, sometimes resorting to unethical practices such as introducing fake papers to increase publication counts. This scenario shows how the drive for quantity in academic outputs can detract from quality, prompting some researchers to inflate their CVs with meaningless publications to meet institutional demands. It has sparked discussions on reforming these pressures, emphasizing quality and integrity over sheer volume .

Publishers sometimes claim stringent peer-review processes as justification for accepting SCIgen papers, such as BEIESP insisting their review structure is rigorous despite evidence to the contrary. Some authors have used SCIgen tools intentionally to test the limits of peer review or to inflate citation metrics and CVs, raising ethical issues about the sincerity of contributions to research and the manipulation of academic records for personal gain. These practices highlight problems in transparency, integrity, and the scholarly responsibility of both publishers and researchers .

Identified SCIgen papers show a significant regional concentration, with 64% from China and 22% from India. This concentration might reflect higher pressures in these regions to produce publications, possibly driven by academic or policy requirements for research output. It suggests that some systemic issues in publishing standards or institutional expectations in these regions could be contributing to the prevalence of such submissions. This pattern prompts a reevaluation of how different regions enforce quality and ethical standards in research .

You might also like