Full Breakdown
AI Agents Struggle in Freelance Work, Challenging Automation Promises
10/31/2025, 1:08:25 PM
Overview of the Research Findings
Recent research conducted by the Center for AI Safety (CAIS) and Scale AI has revealed that leading artificial intelligence (AI) agents are significantly underperforming in simulated freelance tasks. The study utilized a benchmark called the Remote Labor Index to assess the capabilities of various AI models in completing economically valuable work. The results indicated that no AI agent managed to complete more than 3 percent of the assigned tasks, with the highest performer, Manus from a Chinese startup, achieving only 2.5 percent. Other notable models, including Grok from Elon Musk's xAI and Claude from Anthropic, performed similarly poorly, with OpenAI's GPT-5 completing just 1.7 percent of tasks.
Implications for the Workforce
The findings challenge the prevailing narrative that AI will soon replace human workers en masse. Despite the hype surrounding AI's potential, many companies that have attempted to automate tasks have faced disappointing results. For instance, a study from MIT found that 95 percent of companies that piloted AI initiatives reported no meaningful revenue growth. Additionally, the introduction of AI tools has often resulted in low-quality outputs, requiring significant human intervention to correct errors.
Criticism of AI Capabilities
Critics argue that the current generation of AI lacks essential features that limit its effectiveness in complex tasks. Dan Hendrycks, director of CAIS, highlighted that AI agents do not possess long-term memory storage or the ability to learn continuously from experiences, which are crucial for adapting to job requirements. This has led to a growing sentiment among employers that replacing human workers with AI may not yield the expected productivity gains.
Official Statements & Responses
Bing Lie, the director of research at Scale AI, noted that the ongoing debate about AI and jobs has often been theoretical, but the research provides concrete evidence of AI's limitations. He emphasized that many organizations have had to rehire employees after realizing that AI tools were inadequate for their needs. OpenAI did not provide a comment on the study's findings, which contrasts with their previous claims that models like GPT-5 are approaching human-level capabilities.
Verbatim Quotes
- “I should hope this gives much more accurate impressions as to what’s going on with AI capabilities,” — Dan Hendrycks, Director of CAIS
- “We have debated AI and jobs for years, but most of it has been hypothetical or theoretical,” — Bing Lie, Director of Research at Scale AI
- “They don’t have long-term memory storage and can’t do continual learning from experiences.” — Dan Hendrycks, Director of CAIS
Conflicting Reports & Gaps
While the Remote Labor Index presents a stark view of AI's current capabilities, OpenAI's GDPval benchmark claims that frontier AI models are nearing human abilities across various tasks. This discrepancy raises questions about the validity of different assessments of AI performance and highlights the need for further research to reconcile these conflicting reports.
Conclusion
The research underscores the challenges faced by AI in effectively replacing human workers in freelance roles. As companies increasingly turn to automation, the findings serve as a cautionary reminder of the limitations of current AI technologies and the potential consequences for workforce dynamics.
