Drooid Logo
Back to story perspectives

Full Breakdown

Reliability of AI-Generated Citations in Mental Health Research

11/21/2025, 3:55:47 AM

Overview of the Study

A recent study published in JMIR Mental Health has raised significant concerns regarding the reliability of citations generated by advanced artificial intelligence models, particularly OpenAI's GPT-4o. The research, conducted by a team from Deakin University in Australia, focused on the accuracy of bibliographic citations produced by the AI when prompted on various mental health topics. The findings indicate that nearly two-thirds of the citations generated by GPT-4o were either fabricated or contained errors, highlighting a critical risk for scientific research.

Methodology and Findings

The researchers tasked GPT-4o with generating six literature reviews on three mental health conditions: major depressive disorder, binge eating disorder, and body dysmorphic disorder. These conditions were selected based on their varying levels of public recognition and existing research coverage. Each review was required to include at least 20 citations from peer-reviewed academic sources.

Upon analysis of the 176 citations produced, the researchers categorized them into three groups: fabricated (non-existent sources), real with errors (existing sources with inaccuracies), and fully accurate. The results revealed that 35 out of 176 citations, or nearly one-fifth, were entirely fabricated. Furthermore, of the 141 citations that corresponded to real publications, almost half contained at least one error. The study found that the rate of citation fabrication was closely linked to the prominence of the topic; only 6 percent of citations for major depressive disorder were fabricated, compared to 28 percent for binge eating disorder and 29 percent for body dysmorphic disorder.

Implications for Academic Integrity

The implications of this study are profound for the academic community. Researchers are urged to exercise caution when utilizing AI-generated content, emphasizing the necessity for rigorous human verification of references. The findings suggest that academic journals and institutions may need to establish new standards and tools to ensure the integrity of published research in an era increasingly influenced by AI-assisted writing.

Criticism and Limitations

While the study provides valuable insights, it also acknowledges several limitations. The research is specific to the GPT-4o model and the three mental health topics examined, which may not represent the performance of other AI models or a broader range of subjects. Additionally, the prompts used were straightforward, and variations in output could occur with different prompts.

Future Directions

Future research is recommended to explore a wider array of topics and AI models to determine if the observed patterns persist. This could further inform the academic community on the reliability of AI-generated citations and the necessary precautions to take when integrating AI tools into research practices.

Verbatim Quotes

  • “Researchers using these models are advised to exercise caution and perform rigorous human verification of every reference an AI generates.” — Jake Linardon, Lead Researcher
  • “The findings also suggest that academic journals and institutions may need to develop new standards and tools to safeguard the integrity of published research in an era of AI-assisted writing.” — Hannah K Jarman, Co-Author

This study underscores the importance of maintaining rigorous standards in academic research, particularly as reliance on AI technologies continues to grow.