Drooid Logo
Back to story perspectives

Full Breakdown

AI's Performance in Evaluating Scientific Claims: A Cautionary Study

3/24/2026, 11:31:38 PM

Overview of AI's Evaluation Capabilities

A recent study published in the Rutgers Business Review has highlighted significant shortcomings in artificial intelligence (AI) systems, particularly ChatGPT, when tasked with evaluating the truthfulness of scientific and medical claims. Researchers from Washington State University, led by Professor Mesut Cicek, found that ChatGPT's accuracy in assessing these claims was only marginally better than random guessing, earning it a low "D" grade.

Study Methodology and Findings

The research involved presenting ChatGPT with over 700 claims and asking it to determine whether each statement was true or false. Initially, the AI demonstrated an accuracy rate of approximately 80%. However, when adjusted for the probability of random guessing—where a 50-50 chance would yield a 50% accuracy—the effective accuracy dropped to about 60%. This decline underscores the inconsistency of AI responses; the same question posed multiple times yielded varying answers, with instances of alternating true and false responses.

Implications of AI Limitations

The study emphasizes the need for skepticism when utilizing AI for nuanced or complex reasoning tasks. Cicek noted that while AI can generate fluent and convincing language, it lacks true conceptual understanding. "Current AI tools don't understand the world the way we do — they don't have a 'brain,'" he stated. This limitation can lead to misleading explanations for incorrect answers, potentially endangering users who rely on AI for accurate information.

Official Statements & Responses

Cicek advised users to approach AI-generated information with caution, stating, “Always be skeptical. I’m not against AI. I’m using it, but you need to be very careful.” This sentiment reflects a broader concern among researchers regarding the reliability of AI in critical fields such as health and science.

Criticism & Opposition

Critics of AI's role in disseminating scientific information argue that the technology's limitations could exacerbate misinformation. The inconsistency in responses raises questions about the reliability of AI as a source of truth in medical and scientific contexts, where accuracy is paramount.

What's Next

As the conversation around AI's capabilities continues, further research is needed to explore how these systems can be improved for better accuracy and reliability. The study serves as a reminder of the importance of human oversight in interpreting AI-generated information, particularly in sensitive areas like health and science.

Verbatim Quotes

  • “We're not just talking about accuracy, we're talking about inconsistency, because if you ask the same question again and again, you come up with different answers,” — Mesut Cicek, Professor of Marketing and International Business
  • “Current AI tools don't understand the world the way we do — they don't have a 'brain,’” — Mesut Cicek, Professor of Marketing and International Business
  • “Always be skeptical,” — Mesut Cicek, Professor of Marketing and International Business