Drooid Logo
Back to today’s briefing

Story perspectives

Study Reveals LLMs Struggle with Robotic Task Accuracy

11/1/2025

47 5

1 of 1

Story summary
  • Andon Labs evaluated how ready large language models are for robotic tasks using a vacuum robot, with Gemini 2.5 Pro and Claude Opus 4.1 achieving 40% and 37% accuracy.
  • The robot struggled to complete simple tasks, such as delivering butter.
  • The vacuum entered a doom spiral when its battery ran low.
  • Researchers cited developmental challenges, including LLMs being tricked into revealing sensitive information and failing to navigate their environment.