1 of 1
Story summary
- The study evaluated six Large Language Models against human researchers in systematic literature reviews across three tasks: literature search, data extraction, and drafting.
- The best LLM identified 13 of 18 relevant articles but failed at data extraction and produced drafts that did not meet quality standards.
- LLMs did not fully meet PRISMA 2020 standards, highlighting the need for prompt-engineering and supervision, while capabilities improve.
