Drooid Logo
Back to today’s briefing

Story perspectives

GPT-4 Turbo Tops Language Models but Hits 46% Accuracy

1/20/2025

27 3

1 of 1

Story summary
  • In a groundbreaking study, researchers evaluated three major language models—OpenAI's GPT-4, Meta's Llama, and Google's Gemini—using a new benchmark, Hist-LLM. The results revealed that even the top performer, GPT-4 Turbo, managed only 46% accuracy on complex historical questions, underscoring the urgent need for enhancements and addressing biases in training data.