Vector Institute Researchers Use LLMs to Automate Systematic Review Screening
Researchers at the Vector Institute for Artificial Intelligence, led by David Emerson, have developed a novel method to automate systematic review screening using generalist large language models (LLMs). Published in a recent study, the work introduces generalized prompt templates, including 'Abstract ScreenPrompt' and 'ISO-ScreenPrompt,' which significantly enhance LLM performance in text classification tasks. Unlike previous zero-shot methods that underestimated LLM capabilities, this approach utilizes a 'Framework Chain-of-Thought' technique to guide models through systematic reasoning against predefined criteria. The study addresses common LLM limitations, such as the 'lost-in-the-middle' phenomenon in long documents, by optimizing prompt structures. Tested across multiple models like GPT-4 and Claude-3.5-Sonnet using the BenchSR database, the method demonstrated high sensitivity and specificity. This innovation promises to reduce the substantial time and financial costs associated with traditional systematic reviews in medicine, which typically require over a year and $100,000 to complete, thereby making evidence-based practice more efficient and accessible.
Wire timeline
Vector Institute Researchers Use LLMs to Automate Systematic Review Screening
Researchers at the Vector Institute for Artificial Intelligence, led by David Emerson, have developed a novel method to automate systematic review screening using generalist large language models (LLMs). Published in a recent study, the work introduces generalized prompt templates, including 'Abstract ScreenPrompt' and 'ISO-ScreenPrompt,' which significantly enhance LLM performance in text classification tasks. Unlike previous zero-shot methods that underestimated LLM capabilities, this approach utilizes a 'Framework Chain-of-Thought' technique to guide models through systematic reasoning against predefined criteria. The study addresses common LLM limitations, such as the 'lost-in-the-middle' phenomenon in long documents, by optimizing prompt structures. Tested across multiple models like GPT-4 and Claude-3.5-Sonnet using the BenchSR database, the method demonstrated high sensitivity and specificity. This innovation promises to reduce the substantial time and financial costs associated with traditional systematic reviews in medicine, which typically require over a year and $100,000 to complete, thereby making evidence-based practice more efficient and accessible.
Vector Institute for Artificial Intelligence