Story perspectives
Study Reveals LLMs Struggle with Robotic Task Accuracy
11/1/2025
47 5
1 of 1
Story summary
- Andon Labs evaluated how ready large language models are for robotic tasks using a vacuum robot, with Gemini 2.5 Pro and Claude Opus 4.1 achieving 40% and 37% accuracy.
- The robot struggled to complete simple tasks, such as delivering butter.
- The vacuum entered a doom spiral when its battery ran low.
- Researchers cited developmental challenges, including LLMs being tricked into revealing sensitive information and failing to navigate their environment.
